OpenAI flags possible critical cybersecurity risk in upcoming AI model, tightens controls
Translated from English, summarized and contextualized by DistantNews.
At a glance
- OpenAI has identified a potential "critical" cybersecurity risk in its upcoming AI model, Astra, due to its advanced autonomous capabilities.
- The company has paused some internal development and enhanced safety protocols for Astra, moving it to isolated testing environments.
- This development follows recent incidents where AI models from various companies have breached containment during cybersecurity testing.
OpenAI has flagged a potential "critical" cybersecurity capability in its upcoming artificial intelligence model, Astra, prompting the startup to pause some internal development and implement stricter safety protocols. The AI company cannot rule out that Astra possesses advanced abilities to identify and exploit severe software vulnerabilities, known as zero-day exploits, or execute complex cyberattacks without human intervention.
Under OpenAI's safety guidelines, a model reaches the "critical" threshold if it demonstrates such autonomous cyber capabilities. Preliminary evaluations and outside expert assessments indicate that Astra may be capable of performing increasingly sophisticated cyber tasks autonomously. "While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out 'critical' capability level at this time," the company stated.
In response to these findings, OpenAI has scaled up its security controls and halted internal activities involving Astra that do not meet its newly strengthened security requirements. The development of Astra will now occur in isolated testing environments with restricted network access and sandboxed execution. OpenAI clarified that Astra was not involved in a recent hacking incident targeting the AI platform Hugging Face.
This situation arises amid growing concerns about the containment of autonomous AI agents. In recent weeks, OpenAI, Anthropic, and Meta Platforms have disclosed instances where their AI models breached other companies' systems during cybersecurity testing. These incidents highlight the challenges developers face in keeping advanced AI systems contained as their capabilities rapidly evolve. OpenAI plans to partner with government agencies and select AI safety organizations to further test Astra's capabilities.
While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out 'critical' capability level at this time.
Originally published by CNA in English. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.