AI IQ Launches Cybersecurity Benchmarks as Open Source Model Claims ExploitBench Lead
AI IQ announced the launch of AI IQ Cybersecurity, a comprehensive suite of evaluation metrics, alongside OpenPatcher-S1, which achieved a new best score among open source models on ExploitBench. The benchmark platform measures AI models’ ability to identify and exploit software vulnerabilities. Developed jointly by AI IQ and TrustedRouter, OpenPatcher-S1 is a combination model that leverages multiple open source AI systems to coordinate efforts and outperform individual models running alone.
On a specific task testing exploits for CVE-2024-2887, OpenPatcher-S1 achieved a score of 7 out of 16. This places it ahead of the previous open source best, which scored 3 out of 16, but behind the overall frontier model, Anthropic’s Mythos 5, which achieved a perfect 16 out of 16. AI IQ noted that a new proprietary model named Poseidon has already surpassed OpenPatcher-S1 in initial testing but remains in development. Separately, AI IQ’s mini-swe-agent harness is gaining traction across industry and academia for benchmarking tasks, reportedly matching or exceeding the performance of Claude Code and Codex.
From the sources (3 posts)
@ofirpressOur mini-swe-agent is becoming a standard harness for running benchmarks across industry and academia, because it's easy to run & extend while achieving similar or better performance than Claude Code and Codex on benchmarks.
@ryanesheaAnnouncing AI IQ Cybersecurity: the most comprehensive collection of cybersecurity benchmarks. …AND… OpenPatcher-S1: a combination open source model that achieves a SOTA result on ExploitBench among all open source models. OpenPatcher-S1
@ryaneshea@cremieuxrecueil Been attempting to do better on cybersecurity than existing open source models to give people broader access to patching capabilities. Seems to be working: