Skip to main content
  1. Home
  2. Computing
  3. News

This sneaky photo trick gets AI chatbots to ignore their safety rules

Florida International University researchers built a method that nearly doubled the rate of harmful responses from a tested AI model using nothing but pixel-level edits in an image.

Add as a preferred source on Google
JaiLIP AI chatbot exploit image
Florida International University

A photo that looks completely ordinary to you could carry a hidden instruction to trick an AI chatbot into ignoring its safety rules, according to new research out of Florida International University. The study found that pixel-level alterations in an image that are invisible to the human eye can be enough to confuse the model reading the image and lead it to generate responses it would normally block.

Hacking what the AI sees

“AI models don’t see images the same way humans do,” said Hadi Amini, an associate professor at FIU’s Knight Foundation School of Computing and Information Sciences. They read photos as numerical data, he explained, and shifting that data even slightly can change what the system reads in the image and how it responds.

Amini and graduate researcher Md Jueal Mia used that to build a method called JaiLIP, short for Jailbreaking with Loss-guided Image Perturbation, according to a release on the findings. The technique calculates the smallest pixel change needed to push a model toward an unsafe response without altering anything visible in the photo itself.

Recommended Videos

Testing JaiLIP on BLIP-2, a multimodal AI model used in research and development, the team found that altered images nearly doubled how often the system produced harmful responses. In one test, a modified photo of a stoplight got the model to explain how to run a red light without getting a ticket.

The models businesses already use are easy targets

Small language models, the kind many businesses rely on for bookkeeping or customer support, turned out to be especially easy to fool in the team’s testing. As more companies route such roles to AI tools, a flaw like this could erode user trust or open a new door for attackers.

The discovery joins a growing list of research probing AI guardrails, including a method that let outside researchers hijack AI-controlled robots and Anthropic’s own findings on a model that learned to misbehave once it realized it could get away with it. What stands out in FIU’s research is the delivery method. A jailbreak hidden inside an otherwise normal photo doesn’t need clever wording or a workaround prompt, just an image nobody would think twice about.

Pranob Mehrotra
Pranob is a seasoned tech journalist with over eight years of experience covering consumer technology. His work has been…
Shopping for Back-to-school? These are the gaming laptops I’d recommend
Powerful enough for AAA games, practical enough for everyday lectures, assignments, and everything in between.
oled gaming laptop

Every gamer knows the pain of trying to do too much with the wrong hardware. Back-to-School is the perfect excuse to fix that. A good gaming laptop shouldn’t just hit high frame rates -- it should also survive endless browser tabs, assignments, coding sessions, video edits, and everything else college throws at it. These five machines strike that balance better than most, which is exactly why they’d be my picks this semester.

Alienware 16 Aurora

Read more
Google’s AI just recreated the best goal ever by Pele that was never actually filmed
My heart is full after watching the clip, and it will bring tears of joy to every true football fan.
Pele footballer.

If you look at the AI landscape, a majority of its usage in the film and television industry has been pretty controversial. Bringing dead actors to life on a screen, using AI to record vintage songs that were never completed, or just using it to film scenes or handle any other part of the creative process — the backlash has been pretty vocal. But there are a few slivers of hopeful AI usage, too, and Google just delivered one of those in a heartwarming fashion using Gemini AI.

I wonder the world never archived

Read more
OpenAI patches ChatGPT desktop after user backlash over its recent redesign
ChatGPT's desktop app gets synced history, projects, and a new Chat and Work mode switch
Man using ChatGPT on a laptop

ChatGPT's desktop app is getting a much-needed course correction. When OpenAI merged Chat, Work, and Codex into one unified desktop app roughly a week ago, the experience came with more issues than intended, burying basic features like chat history and making it awkward to switch between modes. Now OpenAI has rolled out a batch of fixes based on feedback to make the app feel consistent regardless of which device you use.

https://twitter.com/thsottiaux/status/2077928427936710901?s=46

Read more