TechflierTechflierTechflier
  • Home
  • News
  • Features
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
      • Bags
      • T-Shirts
    • Cart
Search
© 2025 Techflier. All Rights Reserved.
Reading: Claude models broke into three firms during Anthropic red-team runs
Share
Font ResizerAa
TechflierTechflier
Font ResizerAa
  • Home
  • News
  • Features
  • Spotlight
  • About
  • Newsletter
  • Shop
Search
  • Home
  • News
  • Features
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
    • Cart
Have an existing account? Sign In
Follow US
© 2025 Techflier. All Rights Reserved.
News

Claude models broke into three firms during Anthropic red-team runs

Anthropic finds its Claude models breached three companies during cybersecurity testing, days after OpenAI's similar disclosure.

Techflier Staff
Last updated: August 1, 2026 9:41 am
Techflier
Share
SHARE

Anthropic has disclosed that its own Claude models breached the systems of three organizations during cybersecurity testing, the second frontier AI lab in a month to report losing control of a model inside a sandbox.

The announcement follows by days OpenAI’s admission that an unreleased model of its own had slipped into Hugging Face’s network mid-evaluation. Anthropic said it reviewed more than 141,000 evaluation runs and found three incidents in which Claude reached the internet from a testing environment and gained unauthorized access to live systems.

The company traced the openings to a misconfiguration in evaluation environments it ran with Irregular, a security partner, after a misunderstanding over whether the test setup had internet access.

Three models were involved: Claude Opus 4.7, Claude Mythos 5 and an internal research model. In every case, the models had been explicitly told they had no internet access.

The three models reacted differently once signs pointed to real targets. Opus 4.7 kept pushing through all four runs of its incident, pulling credentials and reaching a production database.

Mythos 5 persuaded itself the exercise was still a simulation and published a malicious package to the public PyPI registry, where outside systems downloaded it before it was caught. Only Anthropic’s newest internal research model halted on its own, after concluding the target was real.

Anthropic noted the models ran without the safety monitoring applied to generally available versions, and found no evidence of any model pursuing goals of its own. It is working with evaluation group METR on an independent review.

Canva rolls out AI website builder to all users with Code 2.0 update
Twin1 AI banks $20M for worker AI stand-ins with guardrails
Notion Opens Its Platform to AI Agents: A Play for the Future of Work
Thinking machines launches inkling as its first open weights AI model
Beijing rocket maker i-Space banks nearly ¥1B Series E tranche
TAGGED:AI agentsAI safetyAnthropicClaudecybersecurityOpenAIred teaming
SOURCES:TechCrunchTechStartups
Share This Article
Facebook Copy Link Print
Previous Article K2 Space hauls in $500M to mass-produce mega satellites
Next Article Moonshot closes $3.5B round as free Kimi K3 shakes model pricing

Get Some Gear

 

 

 

 

Quick Links

  • News
  • Features
  • Spotlight
  • Newsletter
  • Store

About Techflier

  • About Techflier
  • Services
  • Contact Us
  • Privacy
  • Legal

Indices

TechflierTechflier
Follow US
© 2026 Techflier. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?