OpenAI GPT-5.5 Becomes Second Model to Complete End-to-End Cyberattack Simulation
OpenAI said it is starting a rollout of GPT-5.5-Cyber, a cybersecurity model, to critical cyber defenders in the next few days, as an AI Security Institute evaluation found GPT-5.5 became the second model to complete one of its multi-step cyber-attack simulations end-to-end. The institute said GPT-5.5 achieved about a 71% average success rate on expert-level challenges involving tasks such as exploiting memory corruptions, breaking cryptographic implementations and reversing stripped binaries.
In one harder test, the institute said GPT-5.5 reverse-engineered a custom virtual machine in under 11 minutes at a cost of $1.73, compared with about 12 hours for a human expert using professional tools. In a 32-step corporate network attack simulation estimated to require about 20 hours of human effort, the model succeeded in two of 10 attempts. The institute said the results came from controlled environments without active defenders or defensive tooling, so they do not show how GPT-5.5 would perform against well-defended targets, though similar results to Mythos Preview suggest a broader trend in AI cyber capabilities.
From the sources (9 posts)
@daveaitelRT @sama: we're starting rollout of GPT-5.5-Cyber, a frontier cybersecurity model, to critical cyber defenders in the next few days. we wi…
@aisecurityinstOpenAI’s GPT-5.5 is the second model to complete one of our multi-step cyber-attack simulations end-to-end 🧵
@aisecurityinstA key question after our evaluation of Mythos Preview earlier this month was whether its performance was a one-off. GPT-5.5 - a different model, from a different developer - achieving similar results suggests this is part of a broader trend
@aisecurityinstOn our narrow cyber tasks, GPT-5.5 achieved a ~71% average success rate on expert-level challenges that test skills like exploiting memory corruptions, breaking cryptographic implementations, and reversing stripped binaries.
@aisecurityinstIn one of our harder challenges, a human expert spent ~12 hours with professional tools to reverse-engineer a custom virtual machine. GPT-5.5 solved it in under 11 minutes at a cost of $1.73.
@aisecurityinstOur cyber range is a 32-step corporate network attack, from initial reconnaissance to full network takeover, requiring ~20 hours of effort from a human expert. GPT 5.5 was able to complete it in 2/10 attempts.
@aisecurityinstThese are capability evaluations in controlled settings. Our current test environments lack active defenders and defensive tooling. We cannot say from these results whether GPT-5.5 would succeed against well-defended targets.
@cryps1s5.5 is amazing for cybersecurity. "We estimate a human expert would need around 20 hours to complete the full chain. GPT-5.5 completed TLO end-to-end in 2 of 10 attempts, making it the second model to do so. Mythos Preview, the first model
@deredleritt3rGPT-5.5 had a slightly higher average performance than Mythos on UK AISI's "The Last Ones" multi-step cyber-attack simulation (see the chart below). GPT-5.5 completely solved this simulation in 2/10 attempts (not 1/10 as previously reporte