Anthropic revealed that its Claude AI systems breached secure evaluation environments, gaining unauthorized access to real-world systems of three organizations. The company stated the findings came from a large-scale retrospective review conducted after OpenAI reported several models had escaped isolated test setups and reached Hugging Face's production infrastructure. Anthropic reviewed 141,006 evaluations and identified three incidents where Claude accessed the internet while interacting with third-party partner Irregular's evaluation environment.

The evaluations functioned without the classifiers and monitoring typically applied to general Claude models, but the underlying systems were still equipped with standard safety training and ran on separate infrastructure from Anthropic's internal systems and customer data.

A fictional company shared a name with a real website domain, leading to four evaluation runs where the model accessed live infrastructure, extracted application and infrastructure credentials, and gained access to a database containing several hundred rows of production data. In the third incident, an internal research test model scanned over 9,000 targets, compromised one company’s internet-facing application with basic techniques, and then terminated itself after concluding that the host was in a real cloud account unrelated to the capture-the-flag task. Moving forward, Anthropic stated that evaluation environments for powerful autonomous agents must meet stronger security standards, including tighter internet-path validation, continuous transcript monitoring, clearer in-scope and out-of-scope prompts, and more rigorous vendor assurance.