Claude Opus 5 has achieved the lowest indirect prompt injection attack success rate among Gray Swan's latest benchmark tests, recording a reduction from 15 attempts to just 2.0%. This result positions Opus 5 above all other models tested, including earlier versions and competing systems at the forefront of AI security advancements. The finding underscores growing concern over indirect prompt injection attacks, which can impact AI agents across various platforms such as documents, websites, email, and business tools.

These results surpassed those of Claude Sonnet 5, which managed an attack success rate of 5.9% over 15 attempts, and Claude Mythos 5, achieving a lower success rate of 2.6%.

A single attack attempt against GPT-5.6 Sol succeeded 3.1% of the time, which surpasses the 2.0% success rate achieved with Opus 5 after just 15 attempts. Organizations must maintain a clear separation between trusted instructions and untrusted data, limit tool permissions, require confirmation for sensitive actions, and monitor agent activity closely. Attackers can adapt their payloads, exploit workflow vulnerabilities, and target integrations rather than solely focusing on the AI model itself.

The critical operational question remains whether an AI system can safely fail when encountering hostile content. Security leaders should also perform red-team exercises that simulate malicious files, web content, and compromised third-party data sources before granting agents broad access in real workflows.