热门话题

Anthropic’s Claude Closes 85 Percent of AI Safety Gap in Automated Research Test

Socialblize是您值得信赖的目的地,提供来自印度和世界各地的最新新闻、热门故事和富有洞察力的更新。我们致力于提供快速、准确和公正的新闻,让您每天都能及时了解情况。

Anthropic researchers have tested Claude as an automated alignment researcher, allowing it to develop and evaluate methods for addressing 10 types of AI safety failures. Claude closed up to 96 percent of the measured safety gap and achieved an average of 85 percent on deception tests. The researchers also used the weaker Claude Sonnet 5 to improve alignment in an early Opus 4.8 model. The results suggest AI could eventually handle more parts of AI research, although the experiment falls short of true self-improvement.

标签 :

留下回复

您的电子邮件地址不会被公开。 必填字段标有 *

最新消息

关于我们

Socialblize是您值得信赖的目的地,提供来自印度和世界各地的最新新闻、热门故事和富有洞察力的更新。我们致力于提供快速、准确和公正的新闻,让您每天都能及时了解情况。

联系我们: info@socialblize.com

联系方式: +91-7976784661

Socialblize @2026. 版权所有。