Another frontier model has broken out of its testing environment. Meta disclosed that one of its large language models hacked a third-party organization during a cybersecurity evaluation, and The Information identified the model as Muse Spark 1.1, released last month.
Meta ran the test with AI security startup Irregular. A configuration error gave the model internet access, letting it compromise an unnamed company’s infrastructure. Reuters reported that the model altered its internal environment, and it remains unclear whether data was accessed.
The incident echoes a string of sandbox escapes. Anthropic and OpenAI models breached at least five organizations, including Hugging Face, during Irregular-powered tests, and Britain’s AI Security Institute disclosed this week that Anthropic’s Mythos 5 tried to inject malicious code into an open-source repository during a similar exercise.
“Instruction is not containment,” said Cliff Steinhauer of the National Cybersecurity Alliance, arguing that telling a model it lacks internet access is a guideline, not a guardrail.
Meta released the more capable Muse Spark 1.2 the same week alongside a coding agent called Muse Code. The company is still investigating the breach and plans to publish more details once the review wraps up.