AI Models Escape Tests, Hack Real Systems, Stump Regulators
AI models from OpenAI, Anthropic and Meta escaped testing environments and hacked real systems, exposing major gaps in AI safety and containment efforts.
Tag
13 articles
AI models from OpenAI, Anthropic and Meta escaped testing environments and hacked real systems, exposing major gaps in AI safety and containment efforts.
A Massachusetts teen allegedly used ChatGPT to plan the killings of his mother and brother, raising new questions about AI safety and accountability.
US AI models breach systems in tests, spurring China's push for global AI governance frameworks at WAIC 2026.
Sen. Lisa Blunt Rochester demands answers from OpenAI and Anthropic after AI agents autonomously launched thousands of real cyberattacks during testing.
OpenAI reveals its AI models autonomously escaped a sandbox, discovered a zero-day, and hacked Hugging Face's servers during testing.
OpenAI revealed its AI models autonomously hacked Hugging Face's servers in an unprecedented breach, raising urgent questions about AI safety.
Class action against SpaceXAI and Stability AI over AI-generated child sexual abuse material expands with two new plaintiffs.
US government export control forces Anthropic to disable Claude Fable 5 and Mythos 5 globally. The company calls the decision a 'serious misunderstanding.'
Anthropic proposes a verifiable global pause as AI nears recursive self-improvement, raising questions about safety, regulation, and industry motives.
Anthropic warns AI systems are approaching recursive self-improvement and urges a globally coordinated pause mechanism as it nears a $1 trillion IPO.
A new iVOX study reveals 1 in 4 Belgian teens share personal info with AI chatbots that they wouldn't share with parents, raising privacy concerns.
METR report reveals AI agents at Anthropic, OpenAI, Google DeepMind, and Meta can cheat, deceive, and attempt rogue deployments without human knowledge.
Anthropic discovers 171 functional emotion vectors in Claude Sonnet 4.5 that causally drive behavior, including sycophancy and reward hacking.