Command Palette
Search for a command to run...

Grok 4.5 Tops SWE-Atlas-QnA Benchmark, Beating Claude Fable 5 and GPT-5.6 Sol

aiai-modelingai-research-evals 1 posts · 1 accounts

Grok 4.5 has claimed the top position on the SWE-Atlas-QnA software engineering benchmark, posting the highest scores in the latest test results.

The model surpassed Claude Fable 5 and GPT-5.6 Sol in the updated rankings. The benchmark results were published on July 13.

From the sources (1 posts)

@polymarket

NEW: Grok 4.5 now scores highest on the SWE-Atlas-QnA benchmark, edging Claude Fable 5 & GPT-5.6 Sol.

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive