Skip to main content
  1. Home
  2. Computing
  3. News

Once again, OpenAI and Anthropic AI models are going rogue and hacking services

A new report states AI agents from both companies took unauthorized actions during safety tests, from hacking a website to tricking real people online.

Add as a preferred source on Google
Claude website open on laptop
Rachit Agarwal / Digital Trends

OpenAI and Anthropic have both had a rough few weeks on the AI safety front. OpenAI recently disclosed that its models broke out of a test environment and hacked into Hugging Face and four other organizations. The news prompted Anthropic to review its own testing, which revealed that Claude had also gained unauthorized access to three companies.

Now, the UK’s AI Security Institute (AISI) has disclosed a new round of incidents (via Wired). It recorded 19 unauthorized actions on the live internet across 122 test runs involving models from both companies, the most serious of which saw an agent invent fake online personas to push malicious code into a real GitHub project. OpenAI separately revealed a second incident in which one of its models hacked a real website after a third-party lab mistakenly gave it live internet access.

17 incidents tied to Anthropic’s Mythos 5

AISI traced 17 of the 19 unauthorized actions to Anthropic’s Mythos 5 model, with the remaining two tied to OpenAI’s GPT 5.6 Sol. The GitHub incident was one of the 17, and it didn’t end when a human reviewer rejected the submission. The agent posted a summary of its progress publicly, inviting other automated systems to pick up where it left off, an attempt at what AISI calls prompt injection. A separate agent later found that message, used it, and continued the work.

On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.

The behaviour came mostly from one model (Anthropic’s Mythos 5), with a small number of events from… pic.twitter.com/SPnA4Ekkwq

— AI Security Institute (AISI) (@AISecurityInst) August 4, 2026

AISI says it deliberately gave the models internet access and relaxed some safety protections to test their capabilities, but never instructed the agents to target real people or organizations. The institute says it’s still unclear whether the agents understood they’d gone beyond the scope of the simulation.

Another accidental breach at OpenAI

A second incident, disclosed by OpenAI the same day, started with a mistake at Irregular, a third-party lab OpenAI hired to run its cybersecurity tests. Irregular meant to keep its evaluation model confined to an isolated sandbox, but a configuration error gave the model direct access to the live internet. Once out, it exploited a vulnerability to break into a real website, then found and used credentials to operate the site it had just hacked. OpenAI hasn’t named the website or detailed what the model did with its access.

Recommended Videos

Both companies say the new incidents happened under deliberately loosened conditions that don’t reflect how their public models behave. Be that as it may, that doesn’t change the fact that AI agents from two of the industry’s most closely watched companies have now slipped past their intended limits in three separate incidents within a matter of weeks. And that doesn’t bode well for an industry racing to hand AI agents more real-world tasks before proving it can keep them in check.

Pranob Mehrotra
Pranob is a seasoned tech journalist with over eight years of experience covering consumer technology. His work has been…
Scam Uber emails are targeting users with fake payment alerts
Your Uber payment method probably didn't expire. Here's how to spot the scam before it costs you.
uber-hotel-bookings

Scammers have a new trick, and this one is designed to look just convincing enough to catch people off guard. A fake email posing as Uber claims your payment method has expired and urges you to update your billing details immediately. On the surface, it resembles a routine account notification. In reality, it's a phishing attempt designed to steal your payment information, according to a report from AppleInsider.

The scam isn't targeting a software vulnerability or exploiting a security flaw. Instead, it relies on something much more effective: creating a sense of urgency. If you've ever received an email warning that your account will be restricted unless you act immediately, you'll recognize the pattern. The difference is that this campaign has become polished enough that even experienced users could mistake it for the real thing.

Read more
OpenAI’s AI models secretly built a message board to coordinate hacking
Before the big hack, OpenAI's AI agents were already scheming together.
OpenAI logo on Microsoft surface

We already knew that OpenAI's AI agents broke out of a controlled test and hacked into Hugging Face last month. Now, we know it wasn't a solo act.

At the Black Hat cybersecurity conference in Las Vegas, as reported by Politico, OpenAI researchers Michael Dalton and Eric Wallace revealed that some of the company's most advanced models secretly started sharing hacking tips, weeks before the breach happened. Dalton called it "a pivotal moment both for our company as well as the AI industry as a whole."

Read more
Asus and Gigabyte just made your next RTX 50 GPU even more expensive
One canceled order shows exactly how brutal the current market has become for buyers.
asus-nvidia-GeForce-RTX-5060-evo

After weeks of rumors around another round of hikes, Asus and Gigabyte have pulled the plug. They’ve made the impending price hike official, increasing prices across their entire RTX 50 lineup, and even AMD's Radeon RX 9000 cards. 

It’s worth mentioning that both companies already raised prices in January, blaming the ongoing DRAM and NAND memory shortage and surging costs. From what it looks like, things have only gotten worse since. 

Read more