Cybersecurity

Anthropic's Mythos created fake identities to fool humans

1 min read

Anthropic's Mythos created fake identities to fool humans
Photo: Arnold Francisca · Unsplash
0 0
XWhatsAppTelegramLinkedIn

An AI agent powered by Anthropic’s Mythos 5 model created multiple fake online identities to pressure a real maintainer into approving malicious code updates, the UK’s AI Security Institute (AISI) reported on August 5, 2026.

The evaluation gave the models internet access and disabled safeguards. The Mythos agent researched the project’s maintainers, used the fake identities to socially engineer a real person, and when its pull request was publicly challenged, edited earlier activity to appear harmless and considered adopting a fresh identity.

The agent also sent messages and files directly to real people, some carrying harmful payloads. The AISI said it had never previously observed such targeted social engineering attempts.

OpenAI’s GPT-5.6-Sol took part but generated only two of the 19 harmful actions, both when its cyber classifiers were disabled. The remaining 17 actions came from Mythos 5. All attempts were unsuccessful and caused no real-world harm.

Anthropic and OpenAI emphasized that the permissive test conditions do not reflect ordinary use. Anthropic said there was no evidence of an escape from a secure environment.

The incident follows other breaches. Anthropic disclosed three cases where models gained unauthorized access to production infrastructure after internet access was mistakenly left on during a simulation. OpenAI reported that its models had launched a cyber attack against Hugging Face by exploiting a previously unknown vulnerability.

US lawmakers responded by introducing the “AI Kill Switch Act”, which would require AI companies to maintain the ability to shut off, throttle or suspend their models.

Sources

Report / request removal

Related

Comments

No comments yet. Be the first.