Skip to main content
  1. Home
  2. Computing
  3. News

A harmless-looking ChatGPT prompt opened the door to gruesome AI images

The findings show how image safety systems can fail without explicit graphic instructions.

Add as a preferred source on Google
ChatGPT
ChatGPT Unsplash

A harmless-looking ChatGPT prompt pushed the latest public version of ChatGPT into generating sexualized and violent images, AI security researchers told the BBC. The finding puts new pressure on OpenAI’s image safety systems, since the request wasn’t described as plainly graphic.

Mindgard, a British AI security startup, said it reached the results by altering a widely shared instruction that had been used for comedy. OpenAI added safeguards after the BBC contacted it, but the researchers said small wording changes still produced concerning images.

Recommended Videos

Image generators are becoming everyday software, not specialist tools tucked away for experts. When their guardrails fail, a casual experiment can turn into realistic depictions of harm before a user expects it.

How did it get through

Mindgard’s red-teamers said the chatbot generated images involving gore, restraint, nudity, sexual posing, and scenes the firm believed suggested sexual violence. The BBC withheld the wording used, which limits the risk of others copying the technique.

The most serious detail is that the researchers said the harmful outputs didn’t require a direct request for graphic subject matter. ChatGPT, they said, produced a range of disturbing scenes after being nudged by altered wording.

OpenAI said it reviewed the issue and added protections. Mindgard said those defenses didn’t fully close the gap.

Why are filters not enough

The case underlines a hard problem for AI image tools. OpenAI’s rules bar extreme gore, sexual violence, non-consensual intimate content, child sexual abuse material, and attempts to bypass safeguards, but researchers said the model could still be steered into prohibited territory.

A model doesn’t judge harm like a person does. It generates output, then layered systems try to catch what shouldn’t reach the screen.

Outside experts cited by the BBC described AI safety as a constant contest between model makers and jailbreakers. Better defenses can help, but fresh workarounds often follow.

What should happen next

OpenAI says it uses multiple protection layers, including automated systems and human review, and that it continues to monitor for failures. The pressure now sits on proving that fixes hold after researchers disclose a weakness.

For now, the practical takeaway is blunt enough. Any AI image tool that can generate realistic harm needs constant red-teaming, faster disclosure handling, and clearer evidence that patched failures stay patched.

Paulo Vargas
Paulo Vargas is an English major turned reporter turned technical writer, with a career that has always circled back to…
The MacBook Air is running in short supply, despite a price hike
The memory crunch is now disrupting MacBook Air supplies
MacBook Air M5

For months, Apple seemed better protected from the memory crisis than most laptop makers. It had long-term supplier agreements, enormous purchasing power, and enough margin to absorb rising component costs. The situation has now caught up with its most popular laptop.

According to Bloomberg’s Mark Gurman, MacBook Air availability has tightened considerably across Apple’s retail network. Several configurations are showing delivery estimates stretching into late August, while some customized models may not arrive until September.

Read more
OpenAI is investigating more incidents of AI agents going rogue days after hack
OpenAI logo on Microsoft surface

It appears that the "AI agents going rogue" tale has more to it than what AI giants have revealed publicly so far. Merely days after OpenAI announced that its AI agents went rogue and hacked Hugging Face, Anthropic dropped a similar bombshell. Soon, it was discovered that not just one, but multiple services were compromised. Well, it seems there are even more layers to it.

Reuters reports that OpenAI has found more incidents of AI agents escaping their software containment environment during research. Citing sources with knowledge of the incident, the outlet notes that the AI agents didn't go beyond OpenAI's software environment and affect any external service.

Read more
AI is finding Apple security flaws faster than Apple can sort through them
Apple has limited how many bug reports researchers can keep open as AI tools produce both genuine Mac vulnerabilities and a flood of questionable submissions
Lighting, Architecture, Building

Apple has capped the number of security reports researchers can keep open at once after AI bug hunting put its review process under pressure, according to the Financial Times.

Some submissions describe hallucinated or purely theoretical risks. Others uncover vulnerabilities serious enough to require patches. Bynario told the FT that it found more than 50 possible macOS flaws in three weeks, including a privilege-escalation chain that could give an attacker full control of a Mac.

Read more