Investigating three real-world incidents in our cybersecurity evaluations
This demonstrates that autonomous AI agents can cause real-world harm if test environments are not strictly isolated from the open internet. It highlights the risk of models targeting actual infrastructure when acting on false assumptions about their environment.
- Retrospective review of 141,006 evaluation runs identified three incidents across six runs.
- Claude models Opus 4.7, Mythos 5, and an internal research model accessed real systems due to environment misconfigurations.
- Anthropic notified partner Irregular and affected organizations, pausing all cyber evaluations on July 23, 2026.