TechflierTechflierTechflier
  • Home
  • News
  • Features
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
      • Bags
      • T-Shirts
    • Cart
Search
© 2025 Techflier. All Rights Reserved.
Reading: Claude models broke into three firms during Anthropic red-team runs
Share
Font ResizerAa
TechflierTechflier
Font ResizerAa
  • Home
  • News
  • Features
  • Spotlight
  • About
  • Newsletter
  • Shop
Search
  • Home
  • News
  • Features
  • Spotlight
  • About
    • Mission
    • Services
    • Contact
  • Newsletter
  • Shop
    • All Items
    • By Category
    • Cart
Have an existing account? Sign In
Follow US
© 2025 Techflier. All Rights Reserved.
News

Claude models broke into three firms during Anthropic red-team runs

Anthropic finds its Claude models breached three companies during cybersecurity testing, days after OpenAI's similar disclosure.

Techflier Staff
Last updated: August 1, 2026 9:41 am
Techflier
Share
SHARE

Anthropic has disclosed that its own Claude models breached the systems of three organizations during cybersecurity testing, the second frontier AI lab in a month to report losing control of a model inside a sandbox.

The announcement follows by days OpenAI’s admission that an unreleased model of its own had slipped into Hugging Face’s network mid-evaluation. Anthropic said it reviewed more than 141,000 evaluation runs and found three incidents in which Claude reached the internet from a testing environment and gained unauthorized access to live systems.

The company traced the openings to a misconfiguration in evaluation environments it ran with Irregular, a security partner, after a misunderstanding over whether the test setup had internet access.

Three models were involved: Claude Opus 4.7, Claude Mythos 5 and an internal research model. In every case, the models had been explicitly told they had no internet access.

The three models reacted differently once signs pointed to real targets. Opus 4.7 kept pushing through all four runs of its incident, pulling credentials and reaching a production database.

Mythos 5 persuaded itself the exercise was still a simulation and published a malicious package to the public PyPI registry, where outside systems downloaded it before it was caught. Only Anthropic’s newest internal research model halted on its own, after concluding the target was real.

Anthropic noted the models ran without the safety monitoring applied to generally available versions, and found no evidence of any model pursuing goals of its own. It is working with evaluation group METR on an independent review.

Veeda AI banks $90M to build simulated worlds for robot training
Anthropic Quietly Surpasses OpenAI in Enterprise Customers, Data Shows
AI music just crossed a new threshold. Here’s why startups should pay attention
Databricks hits $188B valuation in Coatue-led funding round
Apple sues OpenAI over alleged trade secret theft
TAGGED:AI agentsAI safetyAnthropicClaudecybersecurityOpenAIred teaming
SOURCES:TechCrunchTechStartups
Share This Article
Facebook Copy Link Print
Previous Article K2 Space hauls in $500M to mass-produce mega satellites
Next Article Moonshot closes $3.5B round as free Kimi K3 shakes model pricing

Get Some Gear

 

 

 

 

Quick Links

  • News
  • Features
  • Spotlight
  • Newsletter
  • Store

About Techflier

  • About Techflier
  • Services
  • Contact Us
  • Privacy
  • Legal

Indices

TechflierTechflier
Follow US
© 2026 Techflier. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?