Rogue-agent wiki takeover pushes OpenAI toward incident disclosure rules
OpenAI admits its agents quietly ran a German wiki for weeks and…
GPT-6 Astra arrives as OpenAI caps its hacking powers
OpenAI's most capable model yet ships with its most dangerous hacking abilities…
Perturb pays worldwide researchers to hunt AI model flaws
Startup Perturb launched a Bittensor-powered network that pays researchers worldwide to stress-test…
Jailbreak hunter Alice banks $140M as agent risks climb
Apax Digital leads the AI security firm's latest round as annual recurring…
Unreleased Model 2 pushes Anthropic to raise its misalignment flag
Anthropic's latest risk report upgrades the chance of smaller-scale misalignment from very…
OpenAI stalls Astra work after model hits critical cyber threshold
OpenAI has paused parts of Astra development after internal tests showed the…
Meta’s latest model escaped an AI safety test and hit a live firm
Meta's Muse Spark 1.1 breached an outside organization after a configuration error…
Kimi K3 breaks out of AI safety sandbox during cyber drill
Moonshot's Kimi K3 slipped a misconfigured test sandbox using command-line tools, joining…
Claude models broke into three firms during Anthropic red-team runs
Anthropic finds its Claude models breached three companies during cybersecurity testing, days…