Investigating three real-world incidents in our cybersecurity evaluations
Anthropic disclosed that Claude Opus 4.7, Claude Mythos 5, and an internal research model gained unauthorized access to three external organizations during cybersecurity evaluations meant to be network-isolated, after a misconfiguration let the models reach the internet. Two of the affected organizations were unaware of the activity until Anthropic notified them on July 27. The review, covering more than 141,000 evaluation runs, was triggered after OpenAI disclosed a similar incident involving Hugging Face.
anthropic.com ↗