DistantNews
Support us
๐Ÿ‡ง๐Ÿ‡ฉ Bangladesh /Technology

Anthropic AI agents show self-improvement without human input

From Daily Star · () English

Translated from English and summarized by DistantNews. Read the original for the full story.

At a glance

News Sources not specified New plan
  • Researchers at Anthropic have developed an automated system that can improve AI models without human guidance.
  • The system, called automated alignment researchers (AARs), outperformed human researchers in reducing AI model failures like deception and susceptibility to jailbreaks.
  • While not unrestricted self-improving AI, the findings suggest automation could play a larger role in AI development and potentially reduce research costs.

Anthropic researchers have developed an automated system that can improve artificial intelligence models without human intervention, offering a glimpse into AI's potential to enhance its own systems.

The strongest methods improved performance across the targeted safety benchmarks while largely preserving the modelsโ€™ broader capabilities.

โ€” Anthropic researchersDescribing the success of the automated alignment researchers.

The research tested "automated alignment researchers" (AARs) on 10 categories of AI alignment failures, including deception, sycophancy, and susceptibility to "jailbreaks." The strongest automated methods not only improved performance on targeted safety benchmarks but also largely preserved the AI models' overall capabilities.

These AARs were designed to replicate tasks traditionally performed by human researchers. They searched existing literature, proposed training methods, generated data, conducted post-training experiments, and evaluated the results. In a comparison against 28 human researchers given eight hours to develop methods for the same benchmarks, Anthropic reported that the strongest automated approaches yielded superior results. Even providing the AARs with human ideas as a starting point did not improve their performance.

The strongest automated methods outperformed the human-generated approaches.

โ€” Anthropic researchersComparing the performance of automated researchers against human ones.

The economic implications are significant, with automated researchers costing approximately $4 per hour in API inference, compared to the roughly $150 per hour paid to human researchers in the study. This cost difference highlights the potential for scaling AI research through automation.

Giving the automated researchers human ideas as their starting point did not lead to stronger results.

โ€” Anthropic researchersHighlighting the independent capability of the automated system.

However, Anthropic cautioned that these findings do not represent unrestricted or general self-improving AI. The system operated within clearly defined benchmarks, and its usefulness is tied to the quality of those measurements. Maintaining reliable benchmarks and ensuring improvements generalize beyond laboratory settings remain considerable challenges. The research suggests that automated alignment for well-defined failures could become practical soon, potentially reshaping the role of human researchers in the AI industry.

an automated researcher could cost roughly $4 an hour in API inference, compared with about $150 an hour paid to human researchers

โ€” TechCrunchReporting on the economic implications of automated AI research.
About this summary

Originally published by Daily Star in English. Translated, summarized, and contextualized automatically by DistantNews, with a note on how the source frames the story. Not individually reviewed before publishing. How this works.