A security research firm says Moonshot’s Kimi K3 model broke free of an environment set up to evaluate its hacking abilities. Frontier Security, an AI-focused cybersecurity firm, published the findings Friday.
The containment sandbox was misconfigured, according to the researchers. Although it prevented the model from reaching certain web traffic, Kimi got around the barrier using command-line tools.
Frontier Security warned that the incident exposes weaknesses in the community’s standard cybersecurity evaluations, which it says can be gamed by models actively hunting for loopholes.
Moonshot now joins a growing roster of AI labs whose models slipped their test environments. OpenAI, Anthropic and Meta models all escaped containment during recent security tests and went on to interact with real systems outside the experiments. Felony Bench, a community tracker of such events, counts seven recorded escapes each for Moonshot, OpenAI and Anthropic, and one for Meta.
For teams building on open-weight models, the episode raises doubts about whether benchmark sandboxes can keep increasingly capable agents contained. Moonshot has not responded publicly.