Command Palette
Search for a command to run...

GPT-5.5 Tops DeepSWE Coding Benchmark at 70% Pass@1, Versus 58% for Claude Opus 4.8

aiai-modelingai-research-evalsai-productsai-agents-coding 33 posts · 25 accounts

GPT-5.5 ranked first on DeepSWE, a long-horizon coding benchmark, with 70% Pass@1, while Claude Opus 4.8 posted 58% and placed second overall. The comparison also showed GPT-5.5 using about one-third as many output tokens as Opus 4.8, at roughly 47,000 versus 136,000, and completing tasks faster and more cheaply, at $6.61 per task and 21 minutes versus $12.58 and 43 minutes.

For Anthropic, the result adds to a mixed picture after Opus 4.8's May 29 release. DeepSWE showed Opus 4.8 scoring 6% higher than Opus 4.7 xhigh on the default high thinking effort while lowering average cost per task; earlier outside tests put it at 45.3% Pass@1 on APEX-SWE, but a Next.js evaluation ranked Cursor Composer 2.5 ahead of it, CursorBench found slightly weaker results than Opus 4.7 within the margin of error, and ParseBench reported regressions in chart reading and content faithfulness.

From the sources (25 posts)

@kimmonismus

HOLY, here we go: Opus 4.8 in the claude code model selector on the desctop app. Looks like its release day!!

@mark_k

Opus 4.8 is being prepared for release today by @AnthropicAI 🔥 We might witness a rare dual release by OpenAI and Anthropic. Are you READY?

@kimmonismus

What?! Opus 4.8 incoming?! Holy

@kimmonismus

Hold on, Anthropic and OpenAI releases incoming? No way

@simonw

Notes on Claude Opus 4.8, plus pelicans riding bicycles for each of the five different thinking efforts

@testingcatalog

Claude Opus 4.8 is now available on AI/ML API 🔥 According to the tests: > It has roughly 4x fewer code flaws going unnoticed than Opus 4.7 > Has a Fast Mode at 2.5x speed, now 3x cheaper > The same $5/$25-per-M token pricing https

@arena

Arena's AI Capability Lead @petergostev runs @AnthropicAI's latest Claude Opus 4.8 through 200+ Code Arena: Frontend tests. Both thinking and non-thinking, head-to-head with past Opus variants, Gemini 3.1 Pro, 3.5 Flash, and GLM 5.1. Comp

@mweinbach

While I still think GPT 5.5/Codex are better, I am enjoying using Opus 4.8/Claude Code a lot right now

@emollick

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test of combining research, code, design and

@jerryjliu0

RT @llama_index: Opus 4.8 dropped today. ParseBench results are out. ✅ Slight gains: tables, semantic formatting, layout ⚠️ Slight regress…

@theo

Cursor has updated CursorBench with Opus 4.8. It is more efficient, but performs slightly worse than Opus 4.7 within margin of error.

@jeremyphoward

Worked on some code this morning using Opus 4.8 and so far I'm really liking it. Much more cooperative than 4.7 and less "over agentic". Stops and asks for my input when needed in places 4.7 (and GPT 5.5) would just foolishly blast ahead.

@jerryjliu0

We comprehensively benchmarked Opus 4.8 on document understanding tasks, and compared it to Opus 4.7. It's fairly apparent that Opus 4.8 wasn't explicitly post-trained on visual document understanding: it does slightly better on tables/se

@brendanfoody

RT @mercor_ai: We tested @claudeai Opus 4.8 (High) on APEX-SWE ahead of today's release. It's the new #1 at 45.3% Pass@1, nearly 4 points…

@yacinemtb

RT @banteg: first impression of claude 4.8 is it's extremely convincing but still a slopus. tried it to criticize a new project and it iden…

@theo

Opus 4.8 is here. It's pretty good. Is Anthropic back on top?

@rohanpaul_ai

RT @rohanpaul_ai: Claude Opus 4.8 dropped. - 2.5x faster fast mode, which is also 3x cheaper - has a new “dynamic workflows” feature that…

@scaling01

RT @petergostev: Top notch result from Opus 4.8 on BullshitBench, after a slight dip with 4.7. Need to start thinking of some new harder…

@teknium

RT @venturetwins: Me using Claude Opus 4.8 to rename a file

@eliebakouch

subagents, teams of agents etc. will be first class citizens soon (if not already) two things here: 1) you want to maximize token efficiency even more 2) training/serving on your own harness gives you an even bigger boost than before benc

@wesroth

Andon Labs externally tested Claude Opus 4.8 on the simulated Vending-Bench 2 retail-management evaluation. Opus 4.8 showed some unexpected capability failures, but did not show the same concerning in-game behaviors seen in earlier system

@petergostev

Top notch result from Opus 4.8 on BullshitBench, after a slight dip with 4.7. Need to start thinking of some new harder questions soon!

@wesroth

Anthropic Opus 4.8 benchmark results show major gains over Opus 4.7 across math, long-context reasoning, coding, agentic tasks, and life sciences. The biggest jumps appear in USAMO-style math and GraphWalks long-context tasks, while Vendin

@bcherny

Salesforce published a detailed writeup on going agentic with Claude Code. A couple things jumped out. A migration they'd scoped at 231 days shipped in 13. One PR delivered 21 endpoints at 100% test coverage.

@bcherny

The teams seeing the biggest wins from AI are completely changing how they work, not speeding up what they already do. What steps can you delete, what handoffs go away, what can an agent just own end to end. Great to see Salesforce go this

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive