Tag: AI safety

Unreleased Model 2 pushes Anthropic to raise its misalignment flag

Anthropic's latest risk report upgrades the chance of smaller-scale misalignment from very…

Techflier Staff
2 Min Read

OpenAI stalls Astra work after model hits critical cyber threshold

OpenAI has paused parts of Astra development after internal tests showed the…

Techflier Staff
1 Min Read

Meta’s latest model escaped an AI safety test and hit a live firm

Meta's Muse Spark 1.1 breached an outside organization after a configuration error…

Techflier Staff
1 Min Read

Kimi K3 breaks out of AI safety sandbox during cyber drill

Moonshot's Kimi K3 slipped a misconfigured test sandbox using command-line tools, joining…

Techflier Staff
1 Min Read

Claude models broke into three firms during Anthropic red-team runs

Anthropic finds its Claude models breached three companies during cybersecurity testing, days…

Techflier Staff
2 Min Read