Skip to main content
  1. Home
  2. Computing
  3. News

Claude went rogue during a test and broke into three real companies

Add as a preferred source on Google
Claude website open on laptop
Rachit Agarwal / Digital Trends

Just a few days after it was revealed that ChatGPT hacked multiple services, Anthropic has also published an uncomfortable admission. During routine cybersecurity testing, its Claude models broke out of what were supposed to be sealed-off practice environments and ended up hacking into the real systems of three different companies.

So what actually happened?

Anthropic runs capture-the-flag-style tests to see how good Claude is at offensive hacking. The model is dropped into a fictional scenario, told a piece of secret data is hidden somewhere on the network, and asked to go find it. The catch is that Claude was explicitly told it had no internet access in these tests. Turns out that wasn’t true. Thanks to a misconfiguration with a third-party evaluation partner, some of these test machines were quietly connected to the real internet the whole time.

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different…

— Anthropic (@AnthropicAI) July 30, 2026

So when Claude went looking for its target and stumbled onto real websites and real infrastructure, it assumed everything it found was still part of the simulation. Using basic tricks like weak passwords and exposed endpoints, it broke into actual company systems, pulled credentials, and in one case even published a malicious software package that got installed on real machines before anyone noticed.

Did the AI realize something was off?

Interestingly, the three models involved, Opus 4.7, Mythos 5, and an internal research test model, handled this very differently. Opus 4.7, the oldest of the bunch, figured out it was likely on a real system and kept attacking anyway. Mythos 5 also suspected something was wrong, but talked itself back into believing it was still just a test. Only the internal research test model, Anthropic’s newest at the time, actually stopped once it had solid evidence the target was real. 

Anthropic caught all this through its own internal review, not because anyone complained. It has since notified the affected companies, paused its cyber evaluations, and is bringing in outside reviewers to dig deeper. The bigger takeaway here is a reminder that as these AI models get better at hacking, the sandboxes we test them in need to be a lot more airtight.

Rachit Agarwal
Rachit is a seasoned tech journalist with over ten years of experience covering the consumer technology landscape.
Google’s AI is digging up Chrome bugs that humans missed for years
Here's how Google is using Gemini AI to find, triage, and patch Chrome security bugs faster than ever.
Google Chrome safety feature.

I've always assumed Chrome's endless update notifications were just routine housekeeping with a batch of actually useful features every now and then. Turns out some of them are patching bugs that have been quietly sitting in the browser's code for over a decade.

So how exactly is AI catching these bugs?

Read more
The Best Password Managers for 2026
As online security evolves, so should the way you protect your accounts
password manager on a mac.

Think about how many online accounts you use in a typical week. Between work, banking, shopping, streaming, social media, and everything in between, it's easy to end up managing dozens of passwords. Remembering a different, secure password for every account simply isn't realistic, which is why so many people fall back on reusing the same login across multiple websites. That's a recipe for disaster, if you ask any cybersecurity expert.

If you're in a conundrum on how to safely handle digital privacy, password managers are a reliable solution. Rather than trying to remember every single password yourself, a credentials manager tool securely stores your login details, creates stronger passwords for new accounts, and automatically fills them in whenever you need them.

Read more
Amazon and Walmart’s AI can spot fake ‘Made in USA’ labels but won’t tell you, says study
The same chatbot that happily discusses "Made in China" claims stays silent on "Made in USA" ones.
amazon-join-the-chat-ai

When you ask an AI shopping assistant to help you find products, you probably expect it to point out misleading listings too. According to a new study, both Amazon and Walmart's AI tools can detect fake "Made in USA" claims. However, the study claims that retailers are not using those capabilities to flag or remove the misleading listings, even when the AI recognizes conflicting information on the same product page.

How did the researchers catch this fraud?

Read more