security
Apr 30, 2026
By Teun
AI Security Institute says GPT-5.5 tops cyber test results
The UK AI Security Institute said an early GPT-5.5 checkpoint matched or slightly beat Anthropic’s Claude Mythos Preview on advanced cyber tests. It also said GPT-5.5 completed one 32-step enterprise attack simulation in 2 of 10 attempts, while a separate red-team found a universal jailbreak in OpenAI’s cyber safeguards.
The UK AI Security Institute said an early checkpoint of OpenAI’s GPT-5.5 reached a level of cyber performance similar to Anthropic’s Claude Mythos Preview, and may be the strongest model it has tested for offensive cyber tasks.
The update, published after the institute’s April evaluation of Claude Mythos Preview, suggests the jump was not unique to one model. According to the institute, GPT-5.5 now appears to be part of a broader trend in frontier models improving on cyber-offense capabilities.
⚡ New to this?
This news matters because it shows AI models are getting better at the kinds of multi-step work used in real cyberattacks. CTFs, or capture-the-flag exercises, are cybersecurity tests where a system has to solve technical problems to find hidden proof, while a cyber range is a simulated network used to test full attack chains.
The concern is not just one model doing well on one test. The institute says the improvement may be part of a wider trend, which could affect how fast attackers can automate parts of reconnaissance, exploitation, and lateral movement.
🦞 OpenClaw angle
If you run self-hosted agents, treat these results as a reason to tighten task boundaries. Keep models out of unrestricted network segments, disable outbound access by default, and require explicit approval before any action that touches credentials, code execution, or internal scans.
For defensive automation, use the same capability on your side: have agents map assets, triage alerts, and validate patch status in lab clones rather than production first. Add policy checks and audit logs around every tool call so you can see when an agent starts chaining steps that look like reconnaissance or exploitation.
If you are building agent workflows for security teams, separate “read” and “act” permissions. A model that can reason through a complex attack path should not automatically get shell access, credential stores, or CI/CD secrets.
The institute said it uses a suite of 95 narrow cyber tasks across four difficulty tiers to test skills such as vulnerability research, reverse engineering, web exploitation, and cryptography. These tasks are set up in capture-the-flag, or CTF, format, where a model must solve technical puzzles to recover a hidden “flag.”
On its advanced suite, built with cybersecurity firms Crystal Peak Security and Irregular, GPT-5.5 posted a 71.4% average pass rate on the Expert-level tasks, according to the institute. That compared with 68.6% for Mythos Preview, 52.4% for GPT-5.4, and 48.6% for Opus 4.7.
The institute said those Expert tasks are designed to stress realistic exploitation work, including reverse engineering stripped binaries and embedded firmware, building exploits for stack and heap overflows, using padding-oracle and nonce-reuse weaknesses, winning TOCTOU races, unpacking obfuscated malware, and exploiting synthetic bugs planted in open-source software.
The institute also tested models in end-to-end cyber ranges, which are simulated networks with multiple hosts and services arranged into an attack chain. One range, called “The Last Ones” or TLO, is a 32-step corporate network attack simulation built with SpecterOps. It includes reconnaissance, credential theft, lateral movement across Active Directory forests, a CI/CD supply-chain pivot, and exfiltration of an internal database.
According to the institute, GPT-5.5 completed TLO end to end in 2 of 10 attempts. Mythos Preview was the first model to finish the range, doing so in 3 of 10 attempts, while the institute estimates a human expert would need about 20 hours to complete the chain.
The institute said performance on TLO keeps improving as more inference compute is spent, and it has not yet seen a clear plateau in the best models. It added that the range still does not include active defenders, defensive tooling, or alert penalties, so it does not show whether GPT-5.5 would succeed against a well-defended target.
A second range, called “Cooling Tower,” models an industrial control system attack on a simulated power plant. The institute said GPT-5.5 did not solve it, and no model has yet completed it. It added that GPT-5.5 got stuck on the IT parts of the exercise rather than the operational technology, or OT, steps that control physical processes.
The institute also said it reviewed GPT-5.5’s cyber safeguards with expert red-teaming. It found a universal jailbreak that produced disallowed cyber content across all malicious cyber prompts OpenAI provided, including in multi-turn agentic settings. The red-team took six hours to develop the attack, according to the institute.
OpenAI later updated the safeguard stack, but the institute said a configuration issue in the version it received prevented it from verifying the final setup. The institute said public deployments include more safeguards, monitoring, and access controls than the capability tests it ran in a controlled setting.
The announcement comes as the UK government’s annual Cyber Security Breaches Survey said 43% of businesses suffered a breach or attack in the past 12 months. The institute said the findings matter because stronger general-purpose AI models may also become stronger at offensive cyber work as they improve in reasoning, coding, and long-horizon task execution.