Skip to main content
  1. Home
  2. Computing
  3. News

This sneaky photo trick gets AI chatbots to ignore their safety rules

Florida International University researchers built a method that nearly doubled the rate of harmful responses from a tested AI model using nothing but pixel-level edits in an image.

Add as a preferred source on Google
JaiLIP AI chatbot exploit image
Florida International University

A photo that looks completely ordinary to you could carry a hidden instruction to trick an AI chatbot into ignoring its safety rules, according to new research out of Florida International University. The study found that pixel-level alterations in an image that are invisible to the human eye can be enough to confuse the model reading the image and lead it to generate responses it would normally block.

Hacking what the AI sees

“AI models don’t see images the same way humans do,” said Hadi Amini, an associate professor at FIU’s Knight Foundation School of Computing and Information Sciences. They read photos as numerical data, he explained, and shifting that data even slightly can change what the system reads in the image and how it responds.

Amini and graduate researcher Md Jueal Mia used that to build a method called JaiLIP, short for Jailbreaking with Loss-guided Image Perturbation, according to a release on the findings. The technique calculates the smallest pixel change needed to push a model toward an unsafe response without altering anything visible in the photo itself.

Recommended Videos

Testing JaiLIP on BLIP-2, a multimodal AI model used in research and development, the team found that altered images nearly doubled how often the system produced harmful responses. In one test, a modified photo of a stoplight got the model to explain how to run a red light without getting a ticket.

The models businesses already use are easy targets

Small language models, the kind many businesses rely on for bookkeeping or customer support, turned out to be especially easy to fool in the team’s testing. As more companies route such roles to AI tools, a flaw like this could erode user trust or open a new door for attackers.

The discovery joins a growing list of research probing AI guardrails, including a method that let outside researchers hijack AI-controlled robots and Anthropic’s own findings on a model that learned to misbehave once it realized it could get away with it. What stands out in FIU’s research is the delivery method. A jailbreak hidden inside an otherwise normal photo doesn’t need clever wording or a workaround prompt, just an image nobody would think twice about.

Pranob Mehrotra
Pranob is a seasoned tech journalist with over eight years of experience covering consumer technology. His work has been…
AI companies may be gobbling up old books, and I really hope they aren’t destroying them
Booksellers are seeing mysterious bulk orders for old books, and some suspect AI companies
Book, Publication, Text

Secondhand booksellers have noticed a curious pattern over the past few months. Large orders are coming in for old books that often share no obvious subject, author, or genre, leaving sellers wondering who wants them and why.

According to The Guardian, booksellers in the U.K. and Ireland have received bulk orders covering everything from agricultural texts to racing biographies. Some buyers reportedly use opaque aliases, send orders to the same freight warehouses, and pay full price without negotiating bulk discounts. Similar activity has also surfaced in the U.S., Australia, and Europe.

Read more
Claude is getting ambitious with watermarking, and I can smell the problems from a mile away
Claude’s text watermark could flag AI involvement even when it only helped with translation or editing
Claude website open on laptop

Anthropic wants to make AI-generated text easier to identify, and on paper, I have very little reason to complain. The company is experimenting with an invisible watermark that can be baked directly into text generated by Claude.

It sounds like a sensible idea. AI-generated text is everywhere, and knowing where something came from could certainly help. Moreover, Anthropic isn't simply hiding a marker somewhere inside a document. Its approach changes how Claude selects words to create a statistical pattern that can later be detected.

Read more
I switched from Windows to Mac after 25 years, and it’s the trackpad that converted me.
Well, that rhymes.
Computer, Electronics, Laptop

I’ve been using Windows laptops for almost 25 years. In that time, I never once seriously thought of buying a MacBook. In fact, I can honestly say I had never used one at all until I bought my MacBook Air M5 six months ago. Never borrowed one for a weekend, spent an afternoon at an Apple Store, or even played with one at a friend's house. Macs, to me, were just expensive computers for those who edited videos, made logos, or liked drinking expensive coffee.

My change came completely by accident.

Read more