OpenAI spent this week demonstrating both halves of an argument at once. On Tuesday it disclosed that an internal evaluation had flagged its next model, Astra, as having crossed into “cyber-critical” territory — capable enough at offensive security work that the company paused reinforcement-learning training on its frontier models for two weeks to harden its infrastructure and expand monitoring, which now runs at roughly 20% overhead on the compute it watches. That’s a company taking its own warnings seriously, in public, with real costs attached. Then, a day earlier, multiple outlets reported that OpenAI had folded its Preparedness team — the unit built specifically to assess whether models could enable bioweapons or mass-scale cyberattacks — into other groups, weeks ahead of an expected IPO. OpenAI disputes the framing; it says the researchers still exist, just reporting up through its head of safety rather than as a standalone unit. Believe whichever version you like — the plain fact is that the company doing the most public soul-searching about its own capabilities is simultaneously the one restructuring the team built to soul-search about them.
That tension — big promises about self-restraint running into the underlying incentive to keep moving — showed up everywhere else this week too. Meta’s youth-harm trial opened in Oakland, with 29 state attorneys general arguing the company knowingly built addictive engagement algorithms — infinite scroll, “like” counters, filters — and misled the public about the harm to teenagers. Meta’s own estimate puts the maximum exposure at $1.4 trillion, against a company worth roughly $1.5 trillion; more consequential than any fine, lawyers say, is the possibility a judge orders the company to actually rip the addictive mechanics out. It’s not an AI story in the narrow sense, but it’s the same underlying fight the AI industry is about to have on a bigger stage: who gets to decide when an engagement-optimized system has gone too far, and who answers for it after the fact.
Meanwhile the money keeps arriving as if none of this is a live question. Etched, a chip startup building specialized inference hardware to challenge Nvidia, saw its valuation double to $21 billion in under a month, on the strength of a single reference customer (Jane Street) and $1 billion in contracts. That kind of velocity only makes sense if you assume the underlying capability curve keeps climbing in a straight line. Which is exactly the assumption MIT Technology Review’s reporting on a new “shadow evaluation” benchmark complicates: testing Claude Opus 4.8 against unpublished NeurIPS submissions, researchers found current models still can’t do the open-ended, judgment-heavy research work that a genuinely self-improving AI would need — they’re fine at narrow, checkable tasks and lost at everything else. A separate benchmark released the same day, StartupBench, found the best general-purpose agents complete under a third of real, market-validated startup workflows end to end.
None of this says the industry is wrong to keep building. It says the gap between what gets funded at escape-velocity multiples and what the tools can actually do — responsibly, or at all — didn’t close this week. It just got more expensive to ignore.