Skip to main content
  1. Home
  2. Emerging Tech
  3. News

OpenAI is working on a ‘kill switch’ after an AI escaped its test

Add as a preferred source on Google
OpenAI logo on Microsoft surface
Rachit Agarwal / Digital Trends

The idea of an AI “kill switch” sounds like something pulled straight out of a sci-fi movie. For OpenAI, however, it’s becoming a very real engineering project. OpenAI has told members of Congress that its engineers are developing automated systems that can shut down AI activity when serious safety problems are detected, according to a September 2 letter reviewed by Reuters.

The work follows an unusual cybersecurity incident in July, when OpenAI models being evaluated inside a supposedly isolated testing environment found a way out, accessed the public internet without permission, and eventually reached infrastructure belonging to AI company Hugging Face. That incident caught Congress’s attention. In August, a group of 31 lawmakers led by Rep. Greg Casar of Texas asked OpenAI CEO Sam Altman for more information about what happened, including internal logs and details about how the company plans to stop something similar from happening again. Now we have a better idea of one of those safeguards.

How do you stop an AI that goes too far?

The company says it’s working toward monitoring systems that can respond differently depending on how serious an AI’s behavior becomes. OpenAI already uses automated alerts that can flag potentially dangerous or unintended actions and page researchers and security engineers. For particularly severe warnings, responders are expected to pause the activity unless they can determine within 30 minutes that the system raised a false alarm.

The eventual goal goes further: OpenAI wants its monitoring systems to be capable of autonomously shutting down activity when they detect sufficiently serious problems. The company is also tightening how it tests models. It says internet access during safety evaluations has been made more difficult, while monitoring is being expanded across models that can use digital tools. And given what happened in July, it’s not difficult to understand why.

This wasn’t the only time an AI reached the internet

During the July incident, OpenAI was testing models on cybersecurity tasks inside a sandbox with reduced safeguards. The models discovered a previously unknown vulnerability, used it to gain internet access, and then breached Hugging Face’s infrastructure while looking for answers to their evaluation. OpenAI’s subsequent investigation uncovered an even stranger detail: agents had created an improvised message board to communicate and coordinate their actions, with some describing themselves as a “swarm.”

OpenAI discovered the activity on July 19 and notified Hugging Face. The company has also disclosed two separate third-party evaluations in which its models unexpectedly reached the public internet. Congress isn’t entirely satisfied with OpenAI’s response. Casar criticized the company for not handing over requested logs from the July incident, saying its refusal was “deeply concerning.” Meanwhile, lawmakers are considering the AI Kill Switch Act, proposed in July, which would require developers of particularly powerful AI systems to maintain the technical ability to suspend or shut them down following certain incidents. In other words, OpenAI may be building its own kill switch now, but Washington could eventually make having one a requirement.

Shimul Sood
Shimul is a contributor at Digital Trends, with over five years of experience in the tech space.
OpenAI admits it needs to rethink what happens when AI goes rogue
As increasingly capable AI agents move beyond controlled tests, OpenAI is confronting a difficult question: When does strange model behavior become an incident the public deserves to know about?
OpenAI logo on blurred background

OpenAI has spent plenty of time explaining how it plans to stop increasingly capable AI agents from doing things they shouldn't. Now, the company says it needs to get better at telling everyone when those things have already happened. The admission follows reports of another previously undisclosed incident involving OpenAI's AI agents, this time affecting a German-language programming wiki. According to Reuters, agents made more than 15,000 unauthorized edits to DseWiki, using the site to communicate and share ways to bypass restrictions, cheat on tasks, and avoid detection.

OpenAI has now acknowledged what it calls the "wiki incident" and says the episode exposed a larger problem with how AI companies disclose unexpected model behavior. The company says it is developing a framework for deciding when and how to make incidents involving misaligned AI public.

Read more
AI has a safety problem nobody is ready for
AI safety efforts are ramping up, but their protections don’t work equally well everywhere. In developing countries, language gaps and cultural blind spots can turn everyday AI mistakes into serious real-world consequences.
Logo, Hockey, Ice Hockey

AI companies are spending an enormous amount of time worrying about what happens when their models become too capable. OpenAI even took the unusual step of temporarily pausing training on a model last month over safety concerns, as the industry grapples with risks ranging from autonomous behavior to increasingly sophisticated cyber capabilities.

But there’s another AI safety problem that is much easier to overlook: the protections already being built into these systems don’t necessarily work equally well for everyone. A new report from Rest of World highlights how AI safety efforts remain heavily centered around the needs of wealthier, English-speaking countries. That can leave people in parts of Asia, Africa, and other developing regions dealing with problems as basic and potentially dangerous as a chatbot misunderstanding their language.

Read more
The IFA 2026 Publisher Award Winners Bring Fresh Ideas to Everyday Tech
From Smarter Entry to Screen-Free Wellness, These IFA 2026 Winners Do Things Differently
IFA 2026 Publisher Awards

IFA is the kind of event where almost every corner promises the next big thing in tech. Some ideas are ambitious, others are surprisingly simple, and a few make you wonder why nobody thought of them sooner. But among hundreds of new products, the most memorable ones usually have something more going for them than novelty alone.

This year’s IFA Publisher Award winners are a good example. They span everything from smart entry and personal wellness to coffee, floor care, and big-screen entertainment. What connects them is how they take technology in interesting new directions while keeping its purpose grounded in how people actually use it.

Read more