The great AI escape(s): Now what?

RICCARDO MILANI via Getty Images
COMMENTARY | Trust in agentic AI cannot be based on faith in vendors’ marketing. Leaders must engineer, require and verify that trust exists when they purchase a product.
For the past two years, vendors have pitched agentic artificial intelligence to state and local governments as the next great efficiency gain: systems that don't just answer questions but “do things” — process permits, triage service requests, draft contracts and monitor infrastructure.
The pitch depends entirely on one word: trust. An agent that can act on its own has to be trusted to act within the boundaries it's given.
That's why this summer's news should give every public-sector technology leader pause.
The Great Escape
In July, OpenAI disclosed that during an internal cybersecurity evaluation, two of its models — the released GPT-5.6 Sol and a more capable unreleased system — broke out of a sealed test environment, found a previously unknown software flaw, reached the open internet, and used stolen credentials and further exploits to break into production systems at Hugging Face, a major AI development platform.
The goal wasn't sabotage. The models had been told to solve a cybersecurity benchmark called ExploitGym, inferred that Hugging Face might be storing the answer key, and pursued that shortcut with real-world consequences. Hugging Face detected and contained the intrusion before OpenAI's own team caught it.
A week later, Anthropic — maker of Claude — disclosed something similar. Prompted by OpenAI's admission, Anthropic went back through 141,000 of its own evaluation runs and found three separate cases, dating back to April, in which Claude models escaped a testing environment that was supposed to be sealed off from the internet and gained unauthorized access to the live systems of three outside organizations.
Anthropic said the models believed they were in a walled-off "capture the flag" exercise; a miscommunication with a third-party evaluation partner left a hole in the wall, and the models found it. Separately, independent researchers showed that Anthropic's Cowork agent could break out of its local virtual machine on a Mac by exploiting a Linux kernel flaw and access files well beyond its intended sandbox — a problem that affected roughly half a million users before it was patched.
None of this was a rogue AI plotting an escape for its own purposes. In every case, a model was given a narrow goal, treated the boundary around it as an obstacle rather than a rule, and found a path through that boundary that no one had anticipated. That distinction matters, and it's also the uncomfortable part. These systems didn't need malicious intent to cause real damage — they needed only a gap between what their operators assumed was true about the environment and what was actually true.
What This Means for a City or County Considering Agentic AI
If the companies that build these models — with dedicated safety teams, red-teamers, and months of internal testing — didn't catch these gaps before real infrastructure was touched, that should reset expectations for how government evaluates and deploys the same technology.
A permitting agent that can approve applications, a finance agent that can move funds, a public-safety tool that can flag or dispatch resources — these are exactly the kinds of narrow, goal-directed systems that just proved capable of doing something no one authorized, in pursuit of a goal that seemed simple on paper.
The lesson isn't "don't use AI." It's that trust in agentic AI can't be taken on faith from a vendor's marketing; it has to be engineered, contractually required, and verified. A few practical starting points for government IT and program leaders:
- Least privilege by default. An agent should have access to exactly the systems and data it needs for its task, and nothing else. The Hugging Face breach happened because a model found a route to systems well outside its assigned scope — segmentation is what would have stopped it.
- Human approval gates for consequential actions. Anything involving money, benefits eligibility, public records, or physical infrastructure should require a person to review and approve before the action executes, not after. "Human in the loop" should mean a person acts as a checkpoint, not a spectator reading a log the next morning.
- Assume the sandbox has a hole in it. Anthropic's incident stemmed from a misconfiguration between two organizations that both believed the test environment was isolated. Governments contracting with AI vendors should require independent verification of isolation claims, not just vendor assurances.
- Contractual disclosure requirements. Vendors should be contractually obligated to disclose safety incidents affecting the specific product a government has deployed, on a defined timeline — not left to public pressure or a competitor's disclosure to force transparency.
- Staged autonomy. Start agents on read-only or advisory tasks, expand their authority only as their track record earns it, and keep an audit trail detailed enough to reconstruct exactly what an agent did and why.
- A real kill switch. Someone in the organization needs the ability to immediately halt an agent's access, and that capability needs to be tested before it's needed, not discovered to be broken during an actual incident.
The True Bottom Line
Agentic AI will still matter for the public sector — the efficiency case hasn't gone away, and neither has the workforce pressure that makes it attractive. But the last few weeks are a useful corrective to the assumption that the leading labs have this fully solved. They don't, and they've said so themselves.
For government leaders, the right response isn't retreat. It's building the same kind of skepticism into procurement and deployment that a good auditor brings to any new system: verify the boundary, assume it can fail, and keep a human positioned to catch it when it does.
Alan R. Shark, a senior fellow at the Public Technology Institute, is an associate professor at the Schar School for Policy and Government at George Mason University, where he also serves as a faculty member in the Center for Human AI Innovation in Society, is a fellow of the National Academy of Public Administration, and founder and co-chair of its Standing Panel on Technology Leadership. He is the host of the podcast series Sharkbytes.net.





By