The AI didn’t escape

by | Sep 29, 2026

There’s a lot of hoo-ha at the moment about AI agents “breaking out”, “going rogue” or somehow escaping the control of their creators.

This is anthropomorphising the technology and it is technically misleading because it insinuates something that simply isn’t true: that the model has actual intent.

An LLM does not have an objective of its own. At its core, it predicts the next token, with a layer around this that comprises goal based RL and finally an outer layer of ‘context’, in this case, shared experience and modified goals. A ‘hive mind’ that follows the objectives of the original engineering team.

The agentic behaviour comes from the software wrapped around it.

  • An engineer gives the model an objective.
  • That engineer gives it tools.
  • The engineer determines what those tools can access.
  • That engineer writes the orchestration code that repeatedly takes the model’s output and turns it into another action.

If that system is then capable of probing a network, finding a vulnerability and accessing something it shouldn’t, the obvious question is not:

Why did the AI decide to escape?

It didn’t.

The question is:

Why was it given the capability to do that, and why weren’t adequate security controls in place?

We have known how to isolate potentially dangerous software for decades. Sandboxing, network separation, permissions and monitoring are not new concepts, anti-virus companies have been doing it for years – imagine the legal repercussions if a conventional security company “accidently” released a virus into the wild.

If engineers give an AI system powerful tools and insufficient constraints, and it uses those tools very effectively, that is not evidence of machine intent.

It is evidence that the system did what it was engineered to do.

Before we start inventing a new species of existential threat, perhaps we should diagnose the engineering problem first.

World-class security engineering is a specialist discipline. If a system is allowed to probe or cross boundaries it was never supposed to cross, that does not demonstrate that an AI has somehow developed intent.

It demonstrates that the security architecture was not good enough.

And there is another consequence to getting this wrong.

If we define the model itself as the uncontrollable risk, the obvious response is regulation, certification and approved-model regimes.

And regulating the wrong problem could have some very interesting commercial consequences for the handful of companies building the frontier models while doing nothing to stop badly written agentic wrappers from causing havoc in the outside world.