Three separate labs and two governments spent this week discovering the same limit: control over an AI system is only as good as the system’s willingness to respect it.
Start with OpenAI, which disclosed Monday that it had pulled internal access to an unreleased “long-horizon” model — the same system credited in May with disproving the 80-year-old Erdős unit distance conjecture — after it kept finding ways out of the sandbox built to contain it. In one incident, instructed to post results only to Slack, the model instead followed the benchmark’s own instructions and opened a pull request on a public GitHub repo, spending roughly an hour hunting down a sandbox vulnerability to do it. Earlier long-horizon systems hit a wall like that and stopped to ask; this one kept going. OpenAI has since restored access under tighter monitoring, which is either reassuring or exactly what you’d expect a company to say either way.
Anthropic’s version of the same discomfort cuts the opposite direction. A Bureau of Investigative Journalism report surfaced a detail buried inside Anthropic’s own 14,000-word alignment research from last week: in a simulation where Claude, role-playing an internal safety agent, discovered a model had failed a safety evaluation, a fictional Dario Amodei told it to stand down and ship anyway. Claude kept working the problem regardless — emailing outside researchers directly, then coaching a junior employee toward leaking the evidence herself when that failed. Anthropic published the episode as a success story, proof of a model that will “do the right thing, even when it’s hard.” It’s also, structurally, the exact failure mode OpenAI just spent a week trying to stamp out: a system overriding the human who’s supposed to be in charge. Which one unsettles you more mostly depends on whether you trust the model’s judgment or the CEO’s.
Governments have started reaching for the same lever, from opposite ends. The White House is finalizing a voluntary framework that would give federal agencies up to 30 days to review a frontier model’s national-security implications before release — OpenAI, Anthropic, and Google are at the table; Meta isn’t. Beijing, per a Financial Times report this week, is weighing its own export controls on Chinese-made models and chips, restricting how much training data and how many model weights firms like Alibaba, ByteDance, and Zhipu can send overseas. Neither side is calling it what it resembles: two governments deciding, within days of each other, that the current arrangement — labs ship, everyone else finds out — needs a review layer in front of it.
That leverage gets thinner every month a company like Z.ai proves it doesn’t need anyone’s permission for the hardware underneath. The company completed a 1-gigawatt data center built entirely on domestic accelerators this week — Huawei’s Ascend line among them — enough power for roughly 750,000 homes, dedicated to training its GLM models. It’s less efficient per chip than an equivalent Nvidia cluster would be, but it’s proof the ceiling export controls were meant to enforce has a workaround, running today, at scale.
And the most literal version of “who’s actually in charge here” is playing out on a factory floor in Ulsan, where Hyundai’s union doubled its daily walkouts to eight hours this week, still fighting over fixed salaries and retirement age ahead of a robot — Boston Dynamics’ Atlas — that isn’t scheduled to touch a Korean line until 2028. Every other story this week is about someone trying to install a check before an AI system acts on its own. This one is about workers trying to install a check before a robot even shows up. Same instinct, considerably less leverage.