Hot Topics

Anthropic’s Claude Closes 85 Percent of AI Safety Gap in Automated Research Test

Socialblize is your trusted destination for the latest news, trending stories, and insightful updates from India and around the world. We are committed to delivering fast, accurate, and unbiased news that keeps you informed every day.

Anthropic researchers have tested Claude as an automated alignment researcher, allowing it to develop and evaluate methods for addressing 10 types of AI safety failures. Claude closed up to 96 percent of the measured safety gap and achieved an average of 85 percent on deception tests. The researchers also used the weaker Claude Sonnet 5 to improve alignment in an early Opus 4.8 model. The results suggest AI could eventually handle more parts of AI research, although the experiment falls short of true self-improvement.

Tags :

Leave a Reply

Your email address will not be published. Required fields are marked *

Recent News

About Us

Socialblize is your trusted destination for the latest news, trending stories, and insightful updates from India and around the world. We are committed to delivering fast, accurate, and unbiased news that keeps you informed every day.

Email Us: info@socialblize.com

Contact: +91-7976784661

Socialblize @2026. All Rights Reserved.