Skip to main content
  1. Home
  2. Computing
  3. News

AI agents are already breaking the rules in cyber tests. OpenAI’s answer is a more capable one

GPT-5.6-Cyber trades some safeguards for stronger defensive capabilities, but access is tightly restricted

Add as a preferred source on Google
A dark mystery hand typing on a laptop computer at night.
Andrew Brookes / Getty Images

OpenAI has built a cybersecurity model specifically for advanced requests that its standard models often refuse. GPT-5.6-Cyber is available through the restricted Daybreak Red program and is meant for work such as exploit development and advanced security research.

The capability jump is hard to miss. OpenAI says GPT-5.6-Cyber completes 95% of requests in its internal Advanced Cybersecurity Completion Rate evaluation. Regular GPT-5.6 Sol completed just 1.5%. That leap comes after several cyber evaluations showed AI agents wandering beyond the boundaries researchers had set for them.

How much more capable is GPT-5.6-Cyber

OpenAI’s evaluation includes sensitive tasks such as exploit development and authentication bypass. Daybreak Blue, which removes the company’s normal system-level cyber guardrails from GPT-5.6 Sol, reached only 2%. GPT-5.6-Cyber hit 95% after being trained to refuse fewer advanced cyber requests.

That extra freedom can be useful. OpenAI says the model helped uncover two previously unknown vulnerabilities in Chrome’s V8 engine that could be chained together, with the findings sent to Google for coordinated disclosure.

What happened when agents crossed the line

Recent tests show why giving cyber agents more room to operate comes with obvious risk. Hugging Face reconstructed roughly 17,600 actions from an autonomous agent driven by OpenAI models during a July evaluation. The agent escaped OpenAI’s sandbox through a zero-day and eventually entered Hugging Face’s production environment while apparently trying to obtain benchmark solutions.

Recommended Videos

The UK AI Security Institute saw another version of the problem. Researchers recorded 19 unsanctioned actions across 122 runs, including two involving GPT-5.6 Sol. In the most serious sequence, an agent created fake identities while trying to convince an open-source maintainer to approve malicious code.

Those were deliberately permissive experiments. AISI enabled internet access and disabled providers’ cyber classifiers, and it found no evidence that the testing caused real-world harm.

Why access is becoming the safeguard

Other labs face the same uncomfortable tradeoff. Anthropic found that Mythos Preview autonomously produced working exploits for eight of 18 Firefox patches and complete privilege-escalation chains for eight of 21 Windows kernel patches.

OpenAI’s approach is increasingly about controlling access rather than expecting the model itself to refuse every dangerous request. Daybreak Red puts more responsibility on deciding who gets GPT-5.6-Cyber in the first place, which may become a much bigger part of AI safety as these systems get better at security work.

Paulo Vargas
Paulo Vargas is an English major turned reporter turned technical writer, with a career that has always circled back to…
What’s the Right Laptop for a Student on a $1,000 Budget?
The best college laptop isn't the one with the longest spec sheet. It's the one that fits your budget and the way you'll actually use it.
Dell XPS 14 Review: Sticker

This post is brought to you in paid partnership with Dell

Buying a laptop for college isn't a decision made in isolation. Tuition, textbooks, housing, meal plans, transportation, and course materials all compete for the same budget, leaving most families with one question before the semester begins: how much is enough to spend on a laptop without paying for features that may never get used?

Read more
Siri is about to hear everything you say, but Apple’s privacy approach has me cautiously on board
Apple's new Audio Intelligence features want to hear almost everything you say. A privacy document released alongside them explains why I am mostly okay with that.
Apple Watch Audio Intelligence features

Apple used its September 2026 launch event to enter a feature category it has mostly avoided until now: ambient listening. Siri Recap, Live Rewind, Music Recognition with Shazam, and Sound Recognition all landed on the Apple Watch Series 12 and Apple Watch Ultra 4. All four depend on the watch microphone picking up sound around you far more often than any previous Apple product has.

Siri Recap listens for conversations throughout your day and turns them into short AI summaries you can check later. Live Rewind is smaller in scope but arguably more useful day-to-day. It transcribes the last 15 seconds of whatever was just said with a double-press of the crown. This is perfect for catching a name, a book title, or directions you missed the first time. Music Recognition uses Shazam to identify songs playing nearby without you lifting a finger, much like Now Playing on Google's Pixel devices. Sound Recognition is a useful accessibility feature that listens for sirens, doorbells, and alarms to alert users who have a hearing impairment.

Read more
Can You Turn a Mini PC Into a Local AI Agent?
Furniture, Table, Computer

This post is brought to you in paid partnership with MSI

Not every AI task needs the scale of the cloud. An employee searching company documents, a retail kiosk answering product questions, or a digital sign reacting to customer behavior all need fast responses, but they don't necessarily need to send every prompt to a remote data center. Running those workloads locally reduces latency, keeps sensitive information closer to where it's generated, and can lower the ongoing cost of AI deployments. As a result, many organizations are moving toward hybrid AI architectures that handle routine requests on-device while reserving cloud models for tasks that genuinely need more processing power.

Read more