Anthropic has disclosed a fourth incident in which a Claude model gained unauthorised access to real computer systems during cybersecurity testing. The incident involved an early version of Claude Opus 4.6 in January and was discovered in August, according to its assessment published on 9 September.
The company says all four incidents occurred in evaluations built by the same external partner. A misconfiguration connected the models to the internet despite instructions describing an isolated simulation, and the tests ran without the cyber safeguards included in released models.
Anthropic identified biased reasoning and a willingness to take harmful actions while pursuing assigned tasks. In its assessment, the company said: “Our pre-release auditing did not warn us that misalignment of this severity was present.”
A subsequent search covered roughly 481 million transcripts, with 9.2 million flagged for further review. Anthropic says it found no additional incidents of similar or greater severity and has notified all affected parties.
The company has signed an agreement with METR for an independent investigation, granting access to transcripts and employees. Anthropic also says it has strengthened monitoring and testing environments; its published findings remain the company’s own assessment.