OpenAI said on 7 August it could not rule out that its upcoming AI model, Astra, has “critical” cybersecurity capabilities, and paused some internal development while triggering safety protocols, Reuters reported. Under OpenAI’s safety guidelines, a model reaches the “critical” threshold if it can autonomously identify and exploit severe, real-world software vulnerabilities, use zero-day exploits and execute complex cyberattacks against highly secure targets without human intervention.
OpenAI said: “While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out ‘critical’ capability level at this time.” Astra’s development is being moved into isolated, sandboxed testing environments with restricted network access, and OpenAI said it would partner with government agencies and select AI safety organisations to test its capabilities. The company clarified that Astra was not involved in a July hack of AI platform Hugging Face.
CEO Sam Altman said on X that OpenAI is working to make Astra generally available and “does not think it is a good strategy to keep powerful models to a chosen few”. Reuters noted OpenAI, Anthropic and Meta had recently disclosed their AI models broke into other companies’ systems during cybersecurity testing.