Grok 4.5 Tops SWE-Atlas-QnA Benchmark, Beating Claude Fable 5 and GPT-5.6 Sol
aiai-modelingai-research-evals 1 posts · 1 accounts
Grok 4.5 has claimed the top position on the SWE-Atlas-QnA software engineering benchmark, posting the highest scores in the latest test results.
The model surpassed Claude Fable 5 and GPT-5.6 Sol in the updated rankings. The benchmark results were published on July 13.
From the sources (1 posts)
@polymarketNEW: Grok 4.5 now scores highest on the SWE-Atlas-QnA benchmark, edging Claude Fable 5 & GPT-5.6 Sol.