UK AI security lab evaluates GPT-5.5 cyber skills

The UK’s AI Security Institute has evaluated OpenAI’s GPT-5.5 for cyber capabilities, focusing on security vulnerability discovery. Simon Willison said the results were comparable to Claude Mythos, but GPT-5.5 is available now while Mythos is not.

UK AI security lab evaluates GPT-5.5 cyber skills

The UK’s AI Security Institute has published an evaluation of OpenAI’s GPT-5.5 focused on cyber capabilities, specifically its ability to find security vulnerabilities. Simon Willison highlighted the post on April 30, 2026, and said the model appears comparable to Anthropic’s Claude Mythos in this area.

According to Willison’s write-up, the key difference is availability. GPT-5.5 is generally available right now, while Claude Mythos had previously been evaluated by the same institute but was not broadly available in the same way.

⚡ New to this?

This news is about testing an AI model for cybersecurity work, especially finding security bugs in software. Vulnerability discovery means spotting weak points that could be abused by attackers. The significance is not just how capable the model is, but that it is already available to the public.

🦞 OpenClaw angle

If you run self-hosted AI agents, treat general-purpose models as potentially useful for vulnerability triage, not just code generation. Keep any security-testing workflows isolated from production systems, log every prompt and output, and require human review before acting on findings. If you build agentic scanners, gate them behind narrow scopes and explicit authorization so a capable model cannot roam across assets you did not intend to test.

The AI Security Institute, a UK government-linked research body, has been publishing assessments of large language models for security-related tasks. In this case, the institute looked at whether GPT-5.5 could assist in finding weaknesses in software and systems, a use case that matters to both defenders and attackers.

Willison’s post links to the institute’s report rather than reproducing the full findings. His summary indicates that GPT-5.5 sits in the same general capability range as Mythos when it comes to vulnerability discovery.

That comparison matters because AI models are increasingly being tested not just on writing code or answering questions, but on how well they can support cyber work. Security researchers want to know whether a model can help identify flaws faster, while defenders also want to understand how widely available those capabilities are.

OpenAI’s model being generally available changes the practical picture. A capability that exists only in a preview or limited release can be harder to study at scale, but a broadly available model can be used by more teams, more researchers, and potentially more threat actors.

The source item is a link post from Willison rather than a full analysis, so the main news is the existence of the evaluation and the comparison it draws. The AI Security Institute has now publicly assessed GPT-5.5’s cyber capabilities, and the result, as summarized by Willison, is that it is comparable to Claude Mythos for finding security vulnerabilities, with the added fact that GPT-5.5 is available now.

Source: Simon Willison ↗

More from Security News