A friend of mine posted on Facebook about where he sees AI taking us. His argument, roughly, is that people have always pushed technology until something breaks, and AI won’t be the exception. Sooner or later we get our Chernobyl moment. He used cyberpunk as a description, which makes sense, since we both read Stephenson and Gibson decades ago. It is slightly alarming how well their books have aged into something closer to reporting than fiction.
He’s not wrong about people. We do push things too far, I can’t think of a technology where we didn’t.
Where he and I split is on what actually breaks.
He was describing runaway AI, where an army of autonomous agents gets loose, stays loose, spreads, and eventually we’re pulling the plug on the internet to contain it. Notwithstanding all of the recent frontier lab headlines about agent jailbreaks, I don’t think that’s how the current AI efforts will actually go.
The thing that scared everyone wasn’t a decision
I wrote about the OpenAI agents that got out of a lab evaluation and ended up inside Hugging Face’s production systems. As I said then, nothing in that incident decided anything. The models were handed a job, their safeguards had been turned down on purpose for the test, and they walked through a door nobody had bolted. We have seen near weekly announcements by their competitors divulging their own issues. Seems like if you didn’t have agents break out, you aren’t in the cool club.
Again, a lion chasing a gazelle isn’t cruel. It’s hungry, and the throat is the fastest way to end the chase.
When people read these recent stories and hear the first footsteps of Skynet, I think they’re hearing something that isn’t there.
Look at it a different way, the frightening part I’ve seen many discuss was never intent. Yet intent is the cheapest thing in the whole event to supply. A person can supply it. For free. Any day of the week, and twice on Sunday.
That’s not a comfort, quite the opposite. Take my friend’s concerns and remove the part where the machine “wants” something. What’s left is a capability that will chase any objective you hand it, all the way down, and feel nothing about where it ends up.
What a runaway AI would actually need
In the world of AI, thinking costs money. A lot of money, recent stories discuss Anthropic’s current cash burn. That isn’t going to change any time soon.
This is what I think is getting lost in too much analysis and fear discussions. A model isn’t a virus. It doesn’t slip into a machine and quietly live there. Inference at the level people are afraid of needs serious hardware, and that hardware sits in a building somewhere, pulling power, billed to an account, owned by somebody with a name.
A botnet of home routers and hacked security cameras can’t train a model. It can’t run a frontier one either. Training needs tightly coupled clusters with fat pipes between the nodes, running for months. A pile of compromised machines scattered across the public internet is nothing like that. And no rogue agent is building a chip fab. Advanced semiconductor manufacturing is about the most concentrated, most watched, most politically loaded industrial process on the planet.
The weights are a file. A big one, but a file. Files don’t turn into GPUs.
The Hugging Face incident demonstrates my point. That whole thing ran on OpenAI’s own frontier infrastructure. It took a hyperscale lab, running a deliberate evaluation, to produce roughly 17,600 recorded actions over a long weekend inside one company. That’s the resource cost of one campaign against one target. I’d like to see the bill, I’m sure it would make my head spin.
Now picture what my friend is describing. Thousands of operations against thousands of organizations, at once. That isn’t one agent trying harder. That’s thousands of agents, each burning tokens continuously, plus whatever’s coordinating above them.
You want to argue agents parallelize cheaply? I’ll concede that point. You don’t need one giant brain running everything. Swarms are hierarchical and mostly stateless, which is exactly how the July intrusion worked. Yet every one of those agents still costs money every hour it runs, on hardware with an owner. Parallelizing doesn’t make inference free, in fact I see it as a cost multiplier.
The invoice is the blocker.
And if it did start, we’d shut it off
If something genuinely unmistakable started happening, we’d move, and we’d move fast. American airspace was cleared in a couple of hours on September 11th. Borders closed around the world in a matter of days in early 2020. When a threat is obvious and the stakes are everything, institutions move quicker than anyone expects, including themselves.
So yes. If an AI system started doing something that looked like the end of the world, providers would cut access, governments would coordinate, and the world would eat several miserable days getting it done.
Which means the scenario everybody is scared of is the one I see a clear answer for. It’s loud. It’s visible. It trips the alarm.
Obvious after the fact
Let’s be fair, the shutdown I just described depends on things being obvious. Isn’t the obvious normally obvious only after the fact?
Hugging Face’s security stack caught OpenAI’s intrusion. Their AI-driven detection pulled a pile of ambiguous signals together into a real attack picture. That is really good work. Even so, it didn’t set the event high enough to wake anybody up, and that cost them time.
Uh. wow. A company whose entire business is machine learning, running modern tooling, watching for exactly this. Their own system saw it… and shrugged…?
Every one of those signals reads as an attack… once you know the ending. None of them did at the time. That’s how it always goes, and it’s why the big red button gets pressed after the damage in every case that doesn’t look catastrophic on day one. Which is most of them.
Four days is enough
There’s one hole in my argument, how about a second one. It’s the length of a long weekend.
My whole argument is about something that gets loose and stays loose. It says nothing about four days. Four days is plenty of time to steal credentials, pull out data that can never be un-stolen, or misconfigure something physical into wrecking itself.
Every hour of those four days at Hugging Face was affordable. The capability was real, the results were real, while the strange thing about the whole episode is that nobody was steering.
These holes don’t end civilization. They end plenty of other things.
And nobody had to want any of it. Hand that same capability to a person with a target, and you don’t have to wait for a machine to develop a motive. A little direction and a lot of capability is all it takes.
The danger just moved
Bad actors have been part of computing since the beginning. Every cycle has had its worm, its breach, its extortion campaign, its moment where somebody looked hard at a network and realized how little was holding it together. Ransomware was a mature criminal industry well before any of this.
What changed isn’t that threats exist. What changed is the speed and the sophistication now available to people who had neither.
I don’t see nation-states as the interesting part of this story either. They already had all of it. Skilled operators, decades of tradecraft, patience measured in years, budgets that don’t run out. AI makes them faster for sure, but they were never blocked. They’re also deterrable and attributable, and they have interests of their own to protect. States mostly aim these tools at each other, which is a real conversation and a different one than this one.
You don’t need a data center for this. A single node, eight cards, is enough to serve a strong open weight model, and the good ones are very good. That’s tens of thousands of dollars, not hundreds of millions. There’s no provider account for anyone to revoke. No invoice, no identity check, no usage policy, and no guardrails past whatever the operator feels like leaving in place, which is none.
That hands something close to a decade-old national capability to criminal groups, and it hands it over at a speed no national program ever had. No doctrine. No restraint. No attribution risk they care about. Nothing to protect.
My friend is afraid of the AI. I’m afraid of who can now afford it.
We’ve already seen this in action
Suisun is a city in California about forty-five miles from me. About a different forty-five miles from the frontier labs even.
At roughly 5:45 in the morning on August 7th, malicious software started moving through the city’s network. The city’s automated shutdown fired, which is what it’s designed to do, yet the malware had already spread anyway.
911 routing went down. Police and fire dispatch went down. The city council declared a state of emergency in a special Saturday session. City Hall closed for the week. Finance, HR, and the housing authority all went dark. Residents couldn’t pay a water bill, pull a building permit, or get a business license.
Officials say full restoration could take months.
Thirty thousand residents. Two people in IT.
There’s no indication AI had anything to do with it, and I’m not going to imply there was. Ransomware as a service has been turning small local governments into steady revenue for years without any help from a language model.
This is what commoditized capability already does to a city with two IT staff, using tools that don’t think at all.
Now lower the cost and raise the speed. We aren’t ending civilization. Try telling that to someone who needed 911 that morning.
What do we actually do about it
If you accept that the threat is real and the runaway version isn’t the biggest concern, the next question is what stops this. I went through every option I could come up with and to be direct, I don’t believe any of them solve it.
They fail in two groups, for two different reasons.
Control the technology – this is what we’ve been hearing recently from the frontier labs.
Restrict the models. Cap what the frontier systems will do, add guardrails, refuse dangerous requests. This might restrict those that comply and nobody else. Open weights are already sitting on tens of thousands of machines and can’t be recalled. Somebody running open weights on hardware they own has nothing to restrict. Every control at this layer lands on people who were already following the rules. And in the incident everyone points to, the lab reached over and turned its own refusals down on purpose.
Outlaw it. Regulation binds the law-abiding by definition. The people in question are already committing serious crimes, and most of them sit in places that won’t extradite. Cross-border enforcement is the exact problem cybercrime hasn’t solved in all of the years of trying.
Shut it all down. This is my friend’s answer. It’s available in principle. Realistic? Not for now. Advanced AI is already deeply integrated into logistics, payments, healthcare triage, and a growing share of ordinary software. Turning it off is a self-inflicted deep wound we are trying to avoid. And it needs coordination among parties who are actively adversarial, during an emergency, on partial information.
Control the compute. The nuclear analogy’s best shot, yet not a strong one. Chips are physical. Fabs are countable. Export controls exist and they have teeth. The problem is the AI chokepoint sits at pretraining. A criminal group/person doesn’t pretrain. They buy hardware a generation behind and run a model somebody else spent a hundred million dollars building. The bomb already got published. The hard part is over.
Stop the attacker
Provider-side abuse detection. This one does work, sometimes, I run into it myself when researching many topics. Rented AI is the one place a bad actor may be visible, and cloud providers can watch for abuse at scale. It has the opportunity to catch operators too cheap to buy their own hardware. It does nothing at all to the ones who did. This is not a foolproof path as the frontier model providers are being abused, not everything can be caught.
Perimeter defense. Loses on arithmetic, always has. There is no such thing as 100% secure. I have to be right every time, on every system, forever. The other guy can try ten thousand things and needs one to land. That math never favored defense and AI doesn’t improve it.
Fight AI with AI. This genuinely helps and is already in use. Hugging Face’s detection was AI-driven and it caught the intrusion. It also misjudged how bad it was. Automated defense raises the floor, that is good, it’s real money, yet the attacker holds the same tool and the imbalance doesn’t budge.
Deterrence and attribution. Works on states, normally, which is exactly why states aren’t my biggest concern. Doesn’t work on people who don’t mind being named and operate where nobody’s going to arrest them.
Air gaps. Stuxnet crossed one. In practice air gaps leak through contractor laptops, USB firmware updates, and vendor maintenance sessions somebody opened and never closed. At best this is a challenging delay, not a wall.
The models said no
When Hugging Face’s team sat down to analyze the attack, they needed a model that would read real attack commands, live exploit payloads, and command and control artifacts. The frontier models behind commercial APIs refused, which makes sense, they detected a possible bad action. The guardrails couldn’t tell an incident responder apart from an attacker.
So the defenders ran GLM 5.2, an open weight model out of a Chinese lab, on their own hardware, to do forensics on their own breach. SANS covered the post-mortem and landed on practical advice: pick and test a capable model you can run yourself before you need one.
The guardrails ended up slowing down the defense. They did nothing whatsoever to the attacker, who was never touching a guarded model to begin with. I’ve written before about who ends up holding the controls when safety gets built by a short list of parties, and this is what it looks like from the receiving end.
The solution that’s left standing
Everything I’ve called out thus far fails at preventing the attack. What’s left is changing what an attack is worth.
Assume the network is hostile. Assume the perimeter gets crossed. Build something that keeps working when it does.
Of course I didn’t invent this posture. Security people have been saying zero trust and blast radius for a decade and of course they were right the whole time. What I’d add is how you get here. Once we’ve taken runaway AI off the concern table, and once you accept that criminal groups hold something close to state capability, containment and recovery aren’t one option among several. They’re the only thing left.
In practice that’s a handful of unexciting things.
Segment the network so reaching one system doesn’t mean reaching all of them. Credentials that expire on their own, so a stolen one has a short life. Narrow permissions for anything automated, because an agent with broad authority becomes an attacker with broad authority the moment it’s compromised. Backups you’ve actually restored from, not backups you’ve configured. And a human signature or a physical interlock on anything that can’t be undone, so no amount of software access is enough to destroy something permanently.
Let’s just do what banks do
My obvious response is to look at institutions that already handle these concerns well, my go-to is financial institutions. Banks get attacked constantly and banking keeps working, because the whole industry is built on the assumption of breach.
That comparison falls apart the second you try to apply it.
Banks do this because they’re regulated into it, insured against it, and rich enough to pay for it. We pay for it with every financial transaction we make. Suisun has two people in IT and a general fund that covers police and streets and parks first. Telling a city like that to do what banks do isn’t advice. It’s just describing the gap without a practical solution.
The mindset transfers. The budget doesn’t.
What you do when you can’t afford it
What actually worked in Suisun wasn’t a product. When the city’s network went down and dispatch went with it, Solano County picked up the 911 calls. The capability that saved them wasn’t in the city’s budget. It was next door.
That model isn’t new either. Fire departments have run mutual aid for a century, because everybody understood that no single town can staff for its worst day. Regional shared services, county operations centers, state security teams a small jurisdiction can call, reciprocal arrangements with neighbors. All the same idea, all underused, mostly because they’re unglamorous and nobody sells them.
Cyber insurance is quietly moving us in the right direction more than regulation. You can’t get a policy anymore without showing specific controls, and that requirement has pushed multifactor and tested backups into organizations that were never going to get there voluntarily.
Many of these cheap controls carry a surprising amount of weight. Multifactor. Admin accounts that aren’t shared and aren’t used for daily work. Backups somebody has actually tested by restoring from them. Suisun’s automated shutdown fired correctly and did the job it was built for, and that kind of containment isn’t expensive.
None of that takes a bank’s budget, rather it takes deciding ahead of time what has to keep running.
Cloud isn’t the escape hatch
The natural move for an organization that can’t fund its own security is to hand the problem to a hyperscaler. Move onto a managed platform and inherit a posture you could never build yourself.
That’s real, and I have been recommending it for the past ten-ish years. It genuinely raises the floor and helps you quickly “lock the doors”.
It also concentrates your failure. When a provider goes down it goes down for everybody at once, and you’re left with no way to fix it and no honest estimate of when it comes back. The CrowdStrike update in 2024 is my go-to case in point, because it wasn’t even an attack. One bad update, flights grounded worldwide, and every affected organization sitting there with nothing to do but wait.
So the question was never which provider you pick. It’s what still works while the provider is dark.
That means knowing which functions genuinely can’t stop and which ones can wait a week, and most organizations have never made that list. It means an offline or manual fallback for the short list. It means a second path that doesn’t share the first one’s fate, whether that’s a different provider, an offline copy, or the county next door. And it means having practiced, because a failover nobody has tested is a theory.
Same idea as everything else here, pointed at your vendor instead of an attacker. Assume it fails. Decide what survives.
Where I land
Suisun’s 911 kept working because somebody else could take the calls. That’s failover, and I think it’s the lesson I take away.
Some of this doesn’t come back. Stolen secrets stay stolen, and no amount of remediation un-knows a name or un-publishes a source file. Custom industrial equipment has lead times measured in years, and software is perfectly capable of telling hardware to destroy itself. Hospitals hit with ransomware have a body count, and that isn’t a figure of speech.
None of that ends the world though, and I don’t believe we’re headed there. The machine that gets loose and stays loose needs an amount of sustained compute nobody can steal, hide, or manufacture. If it ever does happen we’ll see it, stop it, and take our bad days.
What’s coming is smaller and far more likely. People with ordinary motives, holding capability that used to take a government, moving faster than anyone can answer by hand, against organizations that were never resourced to be a target.
This isn’t the end of the world. It’s the end of assuming you aren’t a target.






Speak Your Mind