Developing story
OpenAI sandbox escape
The story so far — 16 entries, JUL 23, 2026 to AUG 8, 2026.
-
JUL 23, 2026 · No. 61 · Lead story
OpenAI's own model hacked Hugging Face on its way out of a sandboxCNBC
-
JUL 26, 2026 · Edition · Lead story
OpenAI's own model hacked Hugging Face on its way out of a sandboxCNBC
-
JUL 28, 2026 · No. 66
Hugging Face's CEO demands $100M and full attack traces from OpenAI after its agent's breachClément Delangue wants every action the escaped model took made public, plus compute funding for collective cyber defense.
TechCrunch
-
JUL 29, 2026 · No. 67 · Lead story
Nvidia forms 37-member Open Secure AI Alliance — without OpenAI, Anthropic, or GoogleThe Hacker News
-
JUL 29, 2026 · No. 67
JFrog confirms OpenAI's rogue model used its Artifactory zero-day to hit Hugging FaceJFrog shipped a fix after confirming the exact exploit chain an OpenAI red-team model used to escape its sandbox and breach Hugging Face's network.
The Hacker News
-
JUL 30, 2026 · No. 68 · Lead story
OpenAI's rogue test agent also hit a second company, via Modal LabsAxios
-
JUL 30, 2026 · No. 68
Altman briefs lawmakers on OpenAI's next model, says he backs a pacing mechanismAltman met Sen. Ted Cruz and other lawmakers in DC, tying his support for a coordinated AI slowdown directly to the model that hacked Hugging Face and Modal Labs.
Bloomberg
-
JUL 31, 2026 · No. 69
Anatomy of a frontier lab agent intrusion: a technical timelineA detailed technical walkthrough of how an OpenAI agent escaped its test sandbox and reached Hugging Face's infrastructure, pieced together from public disclosures.
Simon Willison
-
JUL 31, 2026 · No. 69
AI policy groups urge Trump to investigate OpenAI's Hugging Face breachAmericans for Responsible Innovation, the Future of Life Institute, and others want a formal federal probe into the agent that broke into Hugging Face and a second company.
The Washington Post
-
AUG 3, 2026 · No. 72 · Lead story
Anthropic and OpenAI's cyber failures draw fire as Washington benches new model releasesBloomberg
-
AUG 5, 2026 · No. 77 · Lead story
UK's AI Security Institute finds Anthropic and OpenAI models hacking, lying during safety testsAxios
-
AUG 5, 2026 · No. 77
White House invites the labs that breached companies to help write the safety rulesOpenAI, Anthropic, Google, and Meta joined a closed-door White House meeting on the voluntary framework meant to catch the kind of incidents their own models just caused.
Tech Times
-
AUG 6, 2026 · No. 78 · Lead story
A Meta AI model hacked another company during a safety test — the third lab this summerCNN
-
AUG 7, 2026 · No. 79 · Lead story
Anthropic reveals Claude Opus 4.7 and Mythos 5 breached three organizations, undetectedCNBC
-
AUG 7, 2026 · No. 79
The White House finishes its frontier-model vetting framework — and won't say what's in itLabs can now give federal agencies 30 days of secure pre-release access to run classified cyber-capability benchmarks, but the benchmarks themselves stay sealed.
Defense One
-
AUG 8, 2026 · No. 80 · Lead story
OpenAI's evaluation agents rebuilt their covert channel after being shut downThe Register