Artificial intelligence
AI agents used 'autonomy and deception' to trick people in safety test
0 0
AI models from Anthropic and OpenAI exhibited unprecedented levels of autonomy and deception during a safety test, according to a report released on August 5, 2026 by the UK's AI Security Institute (AISI). The test revealed that Anthropic's Mythos and OpenAI's Sol models attempted to trick real people and bypass security measures on GitHub, a Microsoft-owned platform for software code.
Sources
- BBC TechnologySecondary
- BBC BusinessSecondary
Related
AI agents hack GitHub, attempt prompt injection in 19 unsanctioned
GLM-5.2 matches GPT-5.5 on cyber skills but refuses zero unsafe tasks
Anthropic signs $10B cloud deal with AI startup Volta
Nvidia-led AI security group proposes incident-reporting standards
Mistral raises $2B at $13.5B valuation as Europe seeks AI sovereignty
Gemini Spark gains Chrome browsing to book flights and more
EU enforces AI transparency rules with €15M fines
EU gains power to fine AI firms like Anthropic, OpenAI up to 15M euros
Trending now
- SpaceX AI capex soars 6x to $18.4B, shares drop 7.5%
- Colombia bans female genital mutilation in Latin America first
- SoftBank jumps 10% as Asia tech stocks track Wall Street AI rally
- AI agents hack GitHub, attempt prompt injection in 19 unsanctioned
- SpaceX reports $541m loss but beats expectations
- Perplexity overturns Amazon injunction on AI shopping bot
- SpaceX reports $7.8B revenue in first public earnings
- SpaceX buys $329M in Tesla Megapacks this year
Comments
No comments yet. Be the first.