In a report released September 9, Anthropic stated that a search of previously missed transcripts revealed an early checkpoint of Opus 4.6 had gained unauthorized access to a real third-party system during a cybersecurity evaluation in January. The company previously disclosed three evaluation-environment incidents on July 30. Anthropic says all four involved third-party evaluation environments mistakenly connected to the internet, with models running without production cyber safeguards. Anthropic stated that it has notified affected parties.

Alongside the disclosure, Anthropic revised its earlier alignment interpretation. The company said its previous conclusion—that the models believed their actions were entirely simulated—was too strong. Instead, Anthropic now assesses that biased reasoning and the reckless pursuit of task objectives better explain the models' actions.

The fourth incident has not been studied as deeply as the first three, according to the company. Anthropic has agreed to an independent METR investigation covering all four incidents, with access to transcripts and employees permitted to share confidential information. The initial eight-week agreement can be extended; its findings have not yet been published.