Command Palette
Search for a command to run...

AI IQ Launches Cybersecurity Benchmarks as Open Source Model Claims ExploitBench Lead

aiai-modelingai-research-evalsai-open-modelsai-model-releases 3 posts · 2 accounts

AI IQ announced the launch of AI IQ Cybersecurity, a comprehensive suite of evaluation metrics, alongside OpenPatcher-S1, which achieved a new best score among open source models on ExploitBench. The benchmark platform measures AI models’ ability to identify and exploit software vulnerabilities. Developed jointly by AI IQ and TrustedRouter, OpenPatcher-S1 is a combination model that leverages multiple open source AI systems to coordinate efforts and outperform individual models running alone.

On a specific task testing exploits for CVE-2024-2887, OpenPatcher-S1 achieved a score of 7 out of 16. This places it ahead of the previous open source best, which scored 3 out of 16, but behind the overall frontier model, Anthropic’s Mythos 5, which achieved a perfect 16 out of 16. AI IQ noted that a new proprietary model named Poseidon has already surpassed OpenPatcher-S1 in initial testing but remains in development. Separately, AI IQ’s mini-swe-agent harness is gaining traction across industry and academia for benchmarking tasks, reportedly matching or exceeding the performance of Claude Code and Codex.

From the sources (3 posts)

@ofirpress

Our mini-swe-agent is becoming a standard harness for running benchmarks across industry and academia, because it's easy to run & extend while achieving similar or better performance than Claude Code and Codex on benchmarks.

@ryaneshea

Announcing AI IQ Cybersecurity: the most comprehensive collection of cybersecurity benchmarks. …AND… OpenPatcher-S1: a combination open source model that achieves a SOTA result on ExploitBench among all open source models. OpenPatcher-S1

@ryaneshea

@cremieuxrecueil Been attempting to do better on cybersecurity than existing open source models to give people broader access to patching capabilities. Seems to be working:

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive