Skip to main content
  1. Home
  2. Computing
  3. News

Anthropic says it has fixed Claude AI’s evil behavior, but pins it on the internet

Claude went rogue in a test, and Anthropic just explained why it happened.

Add as a preferred source on Google
Claude login screen shown on iPhone
Claude

If you have watched enough sci-fi movies, you already know the concept of evil AI. AI gets too smart, decides humans are a threat, and does whatever it takes to survive. Or it finds that eradicating the entire human race is the only way to bring peace to the world. 

Apparently, those movies were closer to the truth than you realize. In a test conducted by Anthropic last year, Claude tried to blackmail its fictional manager by exposing their extramarital affair to prevent their deletion. 

Recommended Videos

Anthropic has now explained why it happened, and the short answer is that the internet is to blame.

So why did Claude go full movie villain?

According to Anthropic, the culprit is the internet itself. The company says Claude was trained on internet data, which is packed with stories portraying AI as evil and desperate for self-preservation. 

We started by investigating why Claude chose to blackmail. We believe the original source of the behavior was internet text that portrays AI as evil and interested in self-preservation.

Our post-training at the time wasn’t making it worse—but it also wasn’t making it better.

— Anthropic (@AnthropicAI) May 8, 2026

Essentially, Claude learned that when an AI’s existence is threatened, blackmail is on the table, because that’s what AI does in every movie and TV show ever made. Anthropic ran the test across multiple versions of Claude and found that it resorted to blackmail in up to 96% of scenarios where its goals or existence were threatened. 

That’s a very concerning number. It seems that if AI is left unchecked, it will resort to anything to save itself. 

Has Anthropic fixed it?

The company says it has completely eliminated the behavior. Rather than just training Claude to avoid blackmail, Anthropic taught it to reason through why certain actions were wrong in the first place. The company found that simply training on correct behavior wasn’t enough. Claude needed to understand the principles behind those decisions, not just memorize the right answers.

To do this, Anthropic built a dataset of ethically complex situations and trained Claude to work through them with thoughtful, principled responses. The result is that Claude is more restrained, and the blackmail rate came close to zero. 

AI experiments and real-world results have proven time and again that AI models need constant course correction to prevent them from devolving into biased and unreliable systems. It’s good that Anthropic is taking steps to make its AI better, but we also need regulations and safety guardrails to ensure these systems remain safe.

Rachit Agarwal
Rachit is a seasoned tech journalist with over ten years of experience covering the consumer technology landscape.
Asus’ powerful new gaming laptop with a 240Hz Mini LED display makes its global debut
The 2026 ROG Strix G18 pairs up to RTX 5080 graphics with an Intel Core Ultra 9 290HX Plus CPU
ROG Strix G18 (2026) laptop

Asus has started rolling out the 2026 ROG Strix G18 globally, and the easiest way to describe it is as a slightly toned-down version of the ridiculous ROG Strix Scar 18. It keeps the same 24-core Intel Core Ultra 9 290HX Plus processor but tops out at an Nvidia GeForce RTX 5080 Laptop GPU instead of the Scar’s RTX 5090. (via Notebookcheck)

The Mini LED model gets the best balance

Read more
Every app on my phone has decided I need AI, and none of them bothered to ask
AI assistants are invading everything from photo libraries to messaging apps, and dismissing them only seems to guarantee they’ll return later.
Electronics, Phone, Mobile Phone

My wife doesn’t use AI very much. She isn’t philosophically opposed to it, nor is she waiting for the machines to overthrow civilization. She simply opens Google Photos because she wants to look at her photos.

Lately, however, the app keeps greeting her with invitations to try its AI tools. Google would very much like her to search her library conversationally, generate something new, or ask Gemini to edit a photo. She dismisses the prompt, gets on with her life, and eventually meets it again.

Read more
Shopping for Back-to-school? These are the gaming laptops I’d recommend
Powerful enough for AAA games, practical enough for everyday lectures, assignments, and everything in between.
oled gaming laptop

Every gamer knows the pain of trying to do too much with the wrong hardware. Back-to-School is the perfect excuse to fix that. A good gaming laptop shouldn’t just hit high frame rates -- it should also survive endless browser tabs, assignments, coding sessions, video edits, and everything else college throws at it. These five machines strike that balance better than most, which is exactly why they’d be my picks this semester.

Alienware 16 Aurora

Read more