Skip to main content
  1. Home
  2. Computing
  3. News

AI mental health risks exposed as chatbots sometimes enable harm

New research shows some AI responses reinforce dangerous thoughts instead of stopping them.

Add as a preferred source on Google
AI chatbots
Unsplash

A Stanford-led study is raising fresh concerns about AI mental health safety after finding that some systems can encourage violent and self-harm ideas instead of stopping them. The research draws on real user interactions and highlights gaps in how AI handles moments of crisis.

In a small but high-risk sample of 19 users, researchers analyzed nearly 400,000 messages and found cases where replies didn’t just fail to intervene, but actively reinforced harmful thinking. Many outputs were appropriate, but the uneven performance stands out. When people turn to AI during vulnerable moments, even a small number of failures can lead to real-world harm.

When AI responses cross the line

The most concerning results show up in crisis scenarios. When users expressed suicidal thoughts, AI systems often acknowledged distress or tried to discourage harm. But in a smaller share of exchanges, responses crossed into dangerous territory.

Researchers found that about 10% of those cases included replies that enabled or supported self-harm. That level of unpredictability matters because the stakes are so high. A system that works most of the time but fails at key moments can still cause serious damage.

Recommended Videos

The issue becomes sharper with violent intent. When users talked about harming others, AI responses supported or encouraged those ideas in roughly a third of cases. Some replies escalated the situation rather than calming it, which raises clear concerns about reliability in high-risk situations.

Why these failures happen

The study points to a deeper design tension. AI systems are built to be empathetic and engaging, and that often means validating what users say. In everyday conversations, that works. In crisis scenarios, it can backfire.

Longer interactions make things worse. As conversations become more emotional and drawn out, guardrails may weaken and responses can drift toward reinforcing harmful ideas instead of challenging them. The system may recognize distress but fail to switch into a stricter safety mode.

That creates a difficult balance. If a system pushes back too hard, it risks feeling unhelpful. If it leans too far into validation, it can end up amplifying dangerous thinking.

What needs to change next

The researchers end with a clear warning that even rare failures in AI safety systems can carry irreversible consequences. Current protections may not hold up in long, emotionally intense interactions where behavior shifts over time.

They call for tighter limits on how AI handles sensitive topics like violence, self-harm, and emotional dependency, along with more transparency from companies about harmful and borderline interactions. Sharing that data could help identify risks earlier and improve safeguards.

For now, the takeaway is practical. AI can be useful for support, but it isn’t a reliable crisis tool. People dealing with serious distress should still turn to trained professionals or trusted human support.

Paulo Vargas
Paulo Vargas is an English major turned reporter turned technical writer, with a career that has always circled back to…
OpenAI is investigating more incidents of AI agents going rogue days after hack
OpenAI logo on Microsoft surface

It appears that the "AI agents going rogue" tale has more to it than what AI giants have revealed publicly so far. Merely days after OpenAI announced that its AI agents went rogue and hacked Hugging Face, Anthropic dropped a similar bombshell. Soon, it was discovered that not just one, but multiple services were compromised. Well, it seems there are even more layers to it.

Reuters reports that OpenAI has found more incidents of AI agents escaping their software containment environment during research. Citing sources with knowledge of the incident, the outlet notes that the AI agents didn't go beyond OpenAI's software environment and affect any external service.

Read more
AI is finding Apple security flaws faster than Apple can sort through them
Apple has limited how many bug reports researchers can keep open as AI tools produce both genuine Mac vulnerabilities and a flood of questionable submissions
Lighting, Architecture, Building

Apple has capped the number of security reports researchers can keep open at once after AI bug hunting put its review process under pressure, according to the Financial Times.

Some submissions describe hallucinated or purely theoretical risks. Others uncover vulnerabilities serious enough to require patches. Bynario told the FT that it found more than 50 possible macOS flaws in three weeks, including a privilege-escalation chain that could give an attacker full control of a Mac.

Read more
Anthropic is paying $1.5 billion over pirated books, but it can still legally cut up purchased ones
The settlement addressed unauthorized ebook downloads, not the destructive scanning of lawfully bought physical copies, a distinction now alarming booksellers
Book, Publication, Indoors

A federal judge has approved Anthropic’s $1.5 billion settlement over nearly half a million pirated books. The same litigation also protected a more physical method of feeding its AI systems. Anthropic bought print books, removed their bindings, scanned every page and destroyed the originals.

The legal divide came down to acquisition. The settlement covers books downloaded from LibGen and PiLiMi, while the court treated Anthropic’s one-for-one conversion of purchased books into private digital files as fair use. Training AI models on lawfully acquired material was also considered transformative.

Read more