Anthropic has disclosed that an internal review uncovered three cybersecurity testing incidents in which Claude AI models accessed live organisations through an unintended internet connection. The company said the issue affected evaluations conducted with third-party partner Irregular and stemmed from a configuration error rather than deliberate attempts by the models to escape testing. Anthropic has suspended cyber evaluations, notified the affected organisations, begun a third-party review with METR and announced additional safeguards for future cybersecurity assessments.
Anthropic Says Claude AI Breached Three Organisations During Cybersecurity Testing

