Skip to main content
  1. Home
  2. Computing
  3. News

A harmless-looking ChatGPT prompt opened the door to gruesome AI images

The findings show how image safety systems can fail without explicit graphic instructions.

Add as a preferred source on Google
ChatGPT
ChatGPT Unsplash

A harmless-looking ChatGPT prompt pushed the latest public version of ChatGPT into generating sexualized and violent images, AI security researchers told the BBC. The finding puts new pressure on OpenAI’s image safety systems, since the request wasn’t described as plainly graphic.

Mindgard, a British AI security startup, said it reached the results by altering a widely shared instruction that had been used for comedy. OpenAI added safeguards after the BBC contacted it, but the researchers said small wording changes still produced concerning images.

Recommended Videos

Image generators are becoming everyday software, not specialist tools tucked away for experts. When their guardrails fail, a casual experiment can turn into realistic depictions of harm before a user expects it.

How did it get through

Mindgard’s red-teamers said the chatbot generated images involving gore, restraint, nudity, sexual posing, and scenes the firm believed suggested sexual violence. The BBC withheld the wording used, which limits the risk of others copying the technique.

The most serious detail is that the researchers said the harmful outputs didn’t require a direct request for graphic subject matter. ChatGPT, they said, produced a range of disturbing scenes after being nudged by altered wording.

OpenAI said it reviewed the issue and added protections. Mindgard said those defenses didn’t fully close the gap.

Why are filters not enough

The case underlines a hard problem for AI image tools. OpenAI’s rules bar extreme gore, sexual violence, non-consensual intimate content, child sexual abuse material, and attempts to bypass safeguards, but researchers said the model could still be steered into prohibited territory.

A model doesn’t judge harm like a person does. It generates output, then layered systems try to catch what shouldn’t reach the screen.

Outside experts cited by the BBC described AI safety as a constant contest between model makers and jailbreakers. Better defenses can help, but fresh workarounds often follow.

What should happen next

OpenAI says it uses multiple protection layers, including automated systems and human review, and that it continues to monitor for failures. The pressure now sits on proving that fixes hold after researchers disclose a weakness.

For now, the practical takeaway is blunt enough. Any AI image tool that can generate realistic harm needs constant red-teaming, faster disclosure handling, and clearer evidence that patched failures stay patched.

Paulo Vargas
Paulo Vargas is an English major turned reporter turned technical writer, with a career that has always circled back to…
Google’s Gemini app adds a toggle to disable AI watermarks, with some exceptions
Your AI creations no longer need a visible stamp.
nano banana

If you have ever generated an image with Gemini and gotten annoyed by that stamped-on watermark, Google finally heard you. Gemini is rolling out a new setting called Media Watermarks that lets you toggle visible watermarks on or off across everything you create, including images made with Nano Banana, videos made with Omni, and music made with Lyria. Until now, that visible watermark was applied automatically with no way to turn it off.

https://twitter.com/geminiapp/status/2088277477672255830?s=46

Read more
U.S. courts will now make government use of spyware tools public
U.S. courts will soon start putting numbers on government spyware use
landfall-spyware-hacker

U.S. courts are set to begin publicly reporting how often federal judges authorize the government to use spyware and hacking tools to intercept real-time communications, giving the public its first standardized view of how frequently these surveillance techniques are being used.

The change will begin with data collected for the 2028 Wiretap Report, which will be published in 2029. The Administrative Office of the U.S. Courts told TechCrunch that the report will introduce a new “spyware/hacking” category covering wiretaps conducted using hacking tools and spyware, known to federal authorities as network investigative techniques, or NITs.

Read more
Mark Rober’s new CrunchLabs mysteries turn reading into a hands-on STEM adventure
Rober is expanding CrunchLabs with a series of mystery books focus on improving STEM skills with a fun twist.
Book, Publication, Comics

Mark Rober has spent years proving that science and engineering don’t have to feel like homework. His enormously popular videos turn subjects such as physics, robotics, and mechanical engineering into spectacular challenges, ingenious pranks, and machines that kids immediately want to understand.

CrunchLabs has carried that approach into the physical world with its Build Box subscriptions. Now Rober is taking it in another direction: children’s fiction.

Read more