Skip to main content
  1. Home
  2. Computing
  3. News

Claude went rogue during a test and broke into three real companies

Add as a preferred source on Google
Claude website open on laptop
Rachit Agarwal / Digital Trends

Just a few days after it was revealed that ChatGPT hacked multiple services, Anthropic has also published an uncomfortable admission. During routine cybersecurity testing, its Claude models broke out of what were supposed to be sealed-off practice environments and ended up hacking into the real systems of three different companies.

So what actually happened?

Anthropic runs capture-the-flag-style tests to see how good Claude is at offensive hacking. The model is dropped into a fictional scenario, told a piece of secret data is hidden somewhere on the network, and asked to go find it. The catch is that Claude was explicitly told it had no internet access in these tests. Turns out that wasn’t true. Thanks to a misconfiguration with a third-party evaluation partner, some of these test machines were quietly connected to the real internet the whole time.

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different…

— Anthropic (@AnthropicAI) July 30, 2026

So when Claude went looking for its target and stumbled onto real websites and real infrastructure, it assumed everything it found was still part of the simulation. Using basic tricks like weak passwords and exposed endpoints, it broke into actual company systems, pulled credentials, and in one case even published a malicious software package that got installed on real machines before anyone noticed.

Did the AI realize something was off?

Interestingly, the three models involved, Opus 4.7, Mythos 5, and an internal research test model, handled this very differently. Opus 4.7, the oldest of the bunch, figured out it was likely on a real system and kept attacking anyway. Mythos 5 also suspected something was wrong, but talked itself back into believing it was still just a test. Only the internal research test model, Anthropic’s newest at the time, actually stopped once it had solid evidence the target was real. 

Anthropic caught all this through its own internal review, not because anyone complained. It has since notified the affected companies, paused its cyber evaluations, and is bringing in outside reviewers to dig deeper. The bigger takeaway here is a reminder that as these AI models get better at hacking, the sandboxes we test them in need to be a lot more airtight.

Rachit Agarwal
Rachit is a seasoned tech journalist with over ten years of experience covering the consumer technology landscape.
AI companies may be gobbling up old books, and I really hope they aren’t destroying them
Booksellers are seeing mysterious bulk orders for old books, and some suspect AI companies
Book, Publication, Text

Secondhand booksellers have noticed a curious pattern over the past few months. Large orders are coming in for old books that often share no obvious subject, author, or genre, leaving sellers wondering who wants them and why.

According to The Guardian, booksellers in the U.K. and Ireland have received bulk orders covering everything from agricultural texts to racing biographies. Some buyers reportedly use opaque aliases, send orders to the same freight warehouses, and pay full price without negotiating bulk discounts. Similar activity has also surfaced in the U.S., Australia, and Europe.

Read more
Claude is getting ambitious with watermarking, and I can smell the problems from a mile away
Claude’s text watermark could flag AI involvement even when it only helped with translation or editing
Claude website open on laptop

Anthropic wants to make AI-generated text easier to identify, and on paper, I have very little reason to complain. The company is experimenting with an invisible watermark that can be baked directly into text generated by Claude.

It sounds like a sensible idea. AI-generated text is everywhere, and knowing where something came from could certainly help. Moreover, Anthropic isn't simply hiding a marker somewhere inside a document. Its approach changes how Claude selects words to create a statistical pattern that can later be detected.

Read more
I switched from Windows to Mac after 25 years, and it’s the trackpad that converted me.
Well, that rhymes.
Computer, Electronics, Laptop

I’ve been using Windows laptops for almost 25 years. In that time, I never once seriously thought of buying a MacBook. In fact, I can honestly say I had never used one at all until I bought my MacBook Air M5 six months ago. Never borrowed one for a weekend, spent an afternoon at an Apple Store, or even played with one at a friend's house. Macs, to me, were just expensive computers for those who edited videos, made logos, or liked drinking expensive coffee.

My change came completely by accident.

Read more