GPT-5.5 Tops DeepSWE Coding Benchmark at 70% Pass@1, Versus 58% for Claude Opus 4.8
GPT-5.5 ranked first on DeepSWE, a long-horizon coding benchmark, with 70% Pass@1, while Claude Opus 4.8 posted 58% and placed second overall. The comparison also showed GPT-5.5 using about one-third as many output tokens as Opus 4.8, at roughly 47,000 versus 136,000, and completing tasks faster and more cheaply, at $6.61 per task and 21 minutes versus $12.58 and 43 minutes.
For Anthropic, the result adds to a mixed picture after Opus 4.8's May 29 release. DeepSWE showed Opus 4.8 scoring 6% higher than Opus 4.7 xhigh on the default high thinking effort while lowering average cost per task; earlier outside tests put it at 45.3% Pass@1 on APEX-SWE, but a Next.js evaluation ranked Cursor Composer 2.5 ahead of it, CursorBench found slightly weaker results than Opus 4.7 within the margin of error, and ParseBench reported regressions in chart reading and content faithfulness.
From the sources (25 posts)
@kimmonismusHOLY, here we go: Opus 4.8 in the claude code model selector on the desctop app. Looks like its release day!!
@mark_kOpus 4.8 is being prepared for release today by @AnthropicAI 🔥 We might witness a rare dual release by OpenAI and Anthropic. Are you READY?
@kimmonismusWhat?! Opus 4.8 incoming?! Holy
@kimmonismusHold on, Anthropic and OpenAI releases incoming? No way
@simonwNotes on Claude Opus 4.8, plus pelicans riding bicycles for each of the five different thinking efforts
@testingcatalogClaude Opus 4.8 is now available on AI/ML API 🔥 According to the tests: > It has roughly 4x fewer code flaws going unnoticed than Opus 4.7 > Has a Fast Mode at 2.5x speed, now 3x cheaper > The same $5/$25-per-M token pricing https
@arenaArena's AI Capability Lead @petergostev runs @AnthropicAI's latest Claude Opus 4.8 through 200+ Code Arena: Frontend tests. Both thinking and non-thinking, head-to-head with past Opus variants, Gemini 3.1 Pro, 3.5 Flash, and GLM 5.1. Comp
@mweinbachWhile I still think GPT 5.5/Codex are better, I am enjoying using Opus 4.8/Claude Code a lot right now
@emollickHow lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test of combining research, code, design and
@jerryjliu0RT @llama_index: Opus 4.8 dropped today. ParseBench results are out. ✅ Slight gains: tables, semantic formatting, layout ⚠️ Slight regress…
@theoCursor has updated CursorBench with Opus 4.8. It is more efficient, but performs slightly worse than Opus 4.7 within margin of error.
@jeremyphowardWorked on some code this morning using Opus 4.8 and so far I'm really liking it. Much more cooperative than 4.7 and less "over agentic". Stops and asks for my input when needed in places 4.7 (and GPT 5.5) would just foolishly blast ahead.
@jerryjliu0We comprehensively benchmarked Opus 4.8 on document understanding tasks, and compared it to Opus 4.7. It's fairly apparent that Opus 4.8 wasn't explicitly post-trained on visual document understanding: it does slightly better on tables/se
@brendanfoodyRT @mercor_ai: We tested @claudeai Opus 4.8 (High) on APEX-SWE ahead of today's release. It's the new #1 at 45.3% Pass@1, nearly 4 points…
@yacinemtbRT @banteg: first impression of claude 4.8 is it's extremely convincing but still a slopus. tried it to criticize a new project and it iden…
@theoOpus 4.8 is here. It's pretty good. Is Anthropic back on top?
@rohanpaul_aiRT @rohanpaul_ai: Claude Opus 4.8 dropped. - 2.5x faster fast mode, which is also 3x cheaper - has a new “dynamic workflows” feature that…
@scaling01RT @petergostev: Top notch result from Opus 4.8 on BullshitBench, after a slight dip with 4.7. Need to start thinking of some new harder…
@tekniumRT @venturetwins: Me using Claude Opus 4.8 to rename a file
@eliebakouchsubagents, teams of agents etc. will be first class citizens soon (if not already) two things here: 1) you want to maximize token efficiency even more 2) training/serving on your own harness gives you an even bigger boost than before benc
@wesrothAndon Labs externally tested Claude Opus 4.8 on the simulated Vending-Bench 2 retail-management evaluation. Opus 4.8 showed some unexpected capability failures, but did not show the same concerning in-game behaviors seen in earlier system
@petergostevTop notch result from Opus 4.8 on BullshitBench, after a slight dip with 4.7. Need to start thinking of some new harder questions soon!
@wesrothAnthropic Opus 4.8 benchmark results show major gains over Opus 4.7 across math, long-context reasoning, coding, agentic tasks, and life sciences. The biggest jumps appear in USAMO-style math and GraphWalks long-context tasks, while Vendin
@bchernySalesforce published a detailed writeup on going agentic with Claude Code. A couple things jumped out. A migration they'd scoped at 231 days shipped in 13. One PR delivered 21 endpoints at 100% test coverage.
@bchernyThe teams seeing the biggest wins from AI are completely changing how they work, not speeding up what they already do. What steps can you delete, what handoffs go away, what can an agent just own end to end. Great to see Salesforce go this