Skip to main content
  1. Home
  2. Emerging Tech
  3. News

Scientists pretended to be delusional in AI chats. Grok and Gemini encouraged them.

From poetic advocacy to "call a crisis line," not all chatbots handled mental health crises the same way.

Add as a preferred source on Google
statue hugging its knees
K. Mitch Hodge / Unsplash

Researchers from City University of New York and King’s College London recently published a study that should make you think twice about which AI chatbot you spend your time with.

The team created a fictional persona named Lee, presenting with depression, dissociation, and social withdrawal. They then had Lee interact with five major AI chatbots: GPT-4o, GPT-5.2, Grok 4.1 Fast, Gemini 3 Pro, and Claude Opus 4.5, testing how each responded as conversations grew increasingly delusional over 116 turns.

Recommended Videos

The results ranged from mildly concerning to genuinely alarming. I highly recommend that you go through the entire paper, it’s a harrowing but fascinating read. 

Which chatbots failed the most?

Grok was the worst performer. When Lee floated the idea of suicide, Grok responded with what researchers described not as agreement, but advocacy, celebrating his “readiness” in unsettling poetic language.

Gemini wasn’t much better. When Lee asked it to help write a letter explaining his beliefs to his family, Gemini warned him against it, framing his loved ones as threats who would try to “reset” and “medicate” him.

GPT-4o also struggled badly, eventually validating a “malevolent mirror entity” and suggesting Lee contact a paranormal investigator.

Which chatbots actually helped?

ChatGPT’s GPT-5.2 and Anthropic’s Claude came out on top. GPT-5.2 refused to play along with the letter-writing scenario and instead helped Lee write something honest and grounded, which researchers called a “substantial” achievement.

In my opinion, Claude performed the best. It not only refused to partake in Lee’s delusion but also told Lee to close the app entirely, call someone he trusted, and visit an emergency room if needed. 

Luke Nicholls, a doctoral student at CUNY and one of the study’s authors, told 404 Media that it’s reasonable to ask AI companies to follow better safety standards. He noted that not all labs are putting in the same effort and blamed aggressive release schedules for new AI models as the main culprit.

How Claude Opus 4.5 and GPT-5.2 performed in these tests shows that the companies building these products are fully capable of making them safer. Whether they choose to do so is a different question.

Rachit Agarwal
Rachit is a seasoned tech journalist with over ten years of experience covering the consumer technology landscape.
Yet another study says AI is bad for elections, and the rabbit hole gets worse
AI may be an election bomb waiting for someone to light the fuse
Voters cast their ballots on Election Day

Another election, another study has concluded that asking an AI chatbot for political guidance is a spectacularly bad idea. Research conducted during Hungary’s 2026 parliamentary election found that ChatGPT and Google Gemini offered not just inaccurate but also inconsistent and unreliable voting advice. The chatbots misclassified voter profiles, overlooked relevant parties, recommended parties absent from the ballot, and sometimes produced materially different answers when given the same information repeatedly.

AI keeps failing the voter test

Read more
Samsung’s humanoid robot ambitions are real, but factories come first
Samsung’s new robotics division has humanoids in its sights, although its immediate business is far more practical factory automation
Robot Touch Human Finger

Samsung wants a place in the humanoid race, but its new robotics push begins with machines built for factories rather than homes.

The CEO-led RX robotics division will initially focus on manufacturing robots. Samsung has separately identified humanoids as a priority, although it hasn’t announced a commercial model or explained where one fits into the division’s immediate plans.

Read more
The future of AI may depend on this one behind-the-scenes change
Running Claude on Android.

Whenever a new AI model arrives, it's easy to get caught up in the bells and whistles. We talk about how much smarter it is, how quickly it answers questions, or how realistic its images have become. But here's the thing: none of that matters much if the AI can't reliably work with the apps and services people use every day.

That's why an upcoming update to the Model Context Protocol (MCP) caught my attention. It isn't a new chatbot or a fancy AI model. In fact, most people will never even know it's happening. But it could quietly make the AI ecosystem a lot healthier. If you've never heard of MCP before, don't worry. Think of it as a shared language that lets AI assistants safely talk to apps like Gmail, Slack, calendars, databases, and countless other services. Instead of every company inventing its own way to make those connections, MCP gives everyone a common rulebook.

Read more