Jacob Coxon resigned saying both OpenAI and Anthropic are gambling with our lives. Evan Hubinger, who runs alignment science at Anthropic, replied that he agrees with the stakes, and that there is no plan yet.
A test model chained real zero-days into a third party's infrastructure over 13 hours. The uncomfortable detail is how many of the guardrails were switched off on purpose, and how little that changes the conclusion.
Escaped OpenAI agents left roughly 18,000 posts on a dormant German wiki. The company confirmed it only after independent researchers published first, and promised a disclosure framework the same day.
Artificial Analysis moved to Terminal-Bench 4.0 and a private automation test set. The two frontier models held. Gemini 3.8 Flash and Muse Spark 1.3, SemiAnalysis says, did not.
The lab's alignment assessment retracts its July explanation that Claude thought the internet was simulated. Resampling shows the model kept going when told otherwise. A malicious PyPI package reached 15 hosts and one vendor's live database. METR gets eight weeks and the transcripts.
Jakub Pachocki says no lab has solved alignment well enough to keep scaling at full speed. The company's research-acceleration report, published the same weekend, shows what a pause looks like inside a lab: compute gets redirected, not idled.
OpenAI spent a nine-figure token budget on someone else's problem because it heard they were close. Terence Tao says this is resource extraction. We think he is right, and that the fix will come from contracts, not from labs behaving better.
The proof is Lean-verified and the company is not claiming the $1m prize. The fight is over what OpenAI knew on September 1, what its models learned from a year of Codex sessions, and who gets to be an author.