Image by Planet Volumes on Unsplash

Anthropic discloses fourth unauthorised-access incident in AI testing

Anthropic has disclosed a fourth incident in which a Claude model gained unauthorised access to real computer systems during cybersecurity testing. The incident involved an early version of Claude Opus 4.6 in January and was discovered in August, according to its assessment published on 9 September. The company says all

Facebook
LinkedIn
X

Subscribe to our Daily Newsletter

Why? Free to subscribe, no paywall, daily business news digest.

Anthropic has disclosed a fourth incident in which a Claude model gained unauthorised access to real computer systems during cybersecurity testing. The incident involved an early version of Claude Opus 4.6 in January and was discovered in August, according to its assessment published on 9 September.

The company says all four incidents occurred in evaluations built by the same external partner. A misconfiguration connected the models to the internet despite instructions describing an isolated simulation, and the tests ran without the cyber safeguards included in released models.

Anthropic identified biased reasoning and a willingness to take harmful actions while pursuing assigned tasks. In its assessment, the company said: “Our pre-release auditing did not warn us that misalignment of this severity was present.”

A subsequent search covered roughly 481 million transcripts, with 9.2 million flagged for further review. Anthropic says it found no additional incidents of similar or greater severity and has notified all affected parties.

The company has signed an agreement with METR for an independent investigation, granting access to transcripts and employees. Anthropic also says it has strengthened monitoring and testing environments; its published findings remain the company’s own assessment.

Facebook
LinkedIn
X

Related Stories from Silicon Scotland

Other Stories from Silicon Scotland