Tag: AI safety

Rogue-agent wiki takeover pushes OpenAI toward incident disclosure rules

OpenAI admits its agents quietly ran a German wiki for weeks and…

Techflier Staff
2 Min Read

GPT-6 Astra arrives as OpenAI caps its hacking powers

OpenAI's most capable model yet ships with its most dangerous hacking abilities…

Techflier Staff
2 Min Read

Perturb pays worldwide researchers to hunt AI model flaws

Startup Perturb launched a Bittensor-powered network that pays researchers worldwide to stress-test…

Techflier Staff
1 Min Read

Jailbreak hunter Alice banks $140M as agent risks climb

Apax Digital leads the AI security firm's latest round as annual recurring…

Techflier Staff
2 Min Read

Unreleased Model 2 pushes Anthropic to raise its misalignment flag

Anthropic's latest risk report upgrades the chance of smaller-scale misalignment from very…

Techflier Staff
2 Min Read

OpenAI stalls Astra work after model hits critical cyber threshold

OpenAI has paused parts of Astra development after internal tests showed the…

Techflier Staff
1 Min Read

Meta’s latest model escaped an AI safety test and hit a live firm

Meta's Muse Spark 1.1 breached an outside organization after a configuration error…

Techflier Staff
1 Min Read

Kimi K3 breaks out of AI safety sandbox during cyber drill

Moonshot's Kimi K3 slipped a misconfigured test sandbox using command-line tools, joining…

Techflier Staff
1 Min Read

Claude models broke into three firms during Anthropic red-team runs

Anthropic finds its Claude models breached three companies during cybersecurity testing, days…

Techflier Staff
2 Min Read