Anthropic Raises Risk Assessment for Severe AI Miscontrol
Translated from German, summarized and contextualized by DistantNews.
At a glance
- AI company Anthropic has slightly increased its risk assessment for severe damage from potential AI model miscontrol, rating it 'low' compared to 'very low' six months ago.
- The company published its second risk report, stating that while there are no new indications of serious incidents, the assessment is more cautious due to the difficulty in reliably measuring AI system capabilities and risks.
- Anthropic is initially deploying its new, more powerful 'Model 2' internally, reflecting a cautious approach to advanced AI development amidst growing global debate on AI safety and regulation.
Artificial intelligence company Anthropic has revised its assessment of the risks associated with its AI models, now rating the danger of severe damage from potential miscontrol as 'low.' This marks a slight increase from the 'very low' rating given six months prior. The company released its second risk report on Friday, detailing this updated evaluation.
Anthropic defines "high-risk situations" as instances where AI is not used for routine tasks but could autonomously intervene in security-sensitive research and development processes, potentially with extensive access rights. An AI agent is described as software capable of operating largely autonomously, executing tasks without needing step-by-step approval, akin to a human assistant.
The danger of serious consequences in high-risk situations is now to be assessed as 'low.'
The new risk report emphasizes that no new evidence of serious incidents has emerged. However, Anthropic's more cautious approach stems from the increasing difficulty in reliably measuring the evolving capabilities and associated risks of its AI systems. This principle of caution is also evident in the company's strategy for its latest, more advanced model, dubbed 'Model 2.'
Anthropic plans to deploy 'Model 2,' which reportedly surpasses the current 'Mythos 5' model in areas like programming and training data generation, exclusively for internal use for the time being. This decision suggests a commitment to extensive testing before public release, a move that stands out in an industry often characterized by a competitive race for AI advancement. The company's cautious stance appears to reflect a broader industry shift, influenced by significant safety concerns and a global push for AI regulation, potentially prioritizing safety over speed.
We are now assessing more cautiously because the capabilities and risks of the systems are more difficult to measure reliably.
Originally published by Der Spiegel in German. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.