Unreleased Model 2 pushes Anthropic to raise its misalignment flag
Anthropic's latest risk report upgrades the chance of smaller-scale misalignment from very…
OpenAI stalls Astra work after model hits critical cyber threshold
OpenAI has paused parts of Astra development after internal tests showed the…
Meta’s latest model escaped an AI safety test and hit a live firm
Meta's Muse Spark 1.1 breached an outside organization after a configuration error…
Kimi K3 breaks out of AI safety sandbox during cyber drill
Moonshot's Kimi K3 slipped a misconfigured test sandbox using command-line tools, joining…
Claude models broke into three firms during Anthropic red-team runs
Anthropic finds its Claude models breached three companies during cybersecurity testing, days…