Command Palette
Search for a command to run...

GPT-5.4 Nano System Matches Gemini 3 Pro at 76.4% on SWE-bench Verified, Paper Finds

aiai-modelingai-research-evalsai-productsai-agents-coding 2 posts · 2 accounts

A new paper reported that a GPT-5.4 nano system using a critic-comparator orchestration loop scored 76.4% on SWE-bench Verified, matching the standalone results cited for Gemini 3 Pro and Claude Opus 4.5 Thinking.

The method selects from eight patch proposals generated by a weak model using execution and proof signals, rather than asking the model to choose the best answer itself. The paper argues that many correct fixes are already present in a weak model's top candidates, making the selector and verification process the key determinant of performance on the benchmark.

From the sources (2 posts)

@dair_ai

NEW paper worth reading. GPT-5.4 nano plus a critic-comparator orchestration loop hits 76.4% on SWE-bench Verified, matching standalone Gemini 3 Pro and Claude Opus 4.5 Thinking. The trick is to select from k=8 weak-model proposals using

@omarsar0

RT @dair_ai: NEW paper worth reading. GPT-5.4 nano plus a critic-comparator orchestration loop hits 76.4% on SWE-bench Verified, matching…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive