Command Palette
Search for a command to run...

150-Line Mini-SWE-Agent Matches or Beats Claude Code, Codex on DeepSWE's 113-Task Coding Benchmark

aiai-modelingai-research-evalsai-productsai-agents-coding 14 posts · 13 accounts

Early results from DeepSWE indicate that mini-swe-agent, a roughly 150-line agent at the core of the evaluation harness, matched or beat Claude Code, Codex and Gemini CLI on the new agentic coding benchmark. DeepSWE spans 113 tasks across 91 repositories in five languages and is intended to show where top coding agents diverge in day-to-day developer workflows, even when public leaderboards make them look close in performance.

Each model is tested with the same system instructions and a single bash tool, without vendor-specific editing primitives. While the prompts are shorter than SWE-Bench Pro, they require 5.5 times more code and touch seven files on average; because the workflow for finding code, reproducing issues, fixing them and verifying results maps directly to the verifier, the benchmark may partly reward adherence to that process rather than raw coding ability alone.

From the sources (14 posts)

@garrytan

This is the new standard for engineering evals

@theo

This is the first code bench that actually aligns with how it feels to use these models coding.

@mweinbach

RT @serenaa_ge: Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look…

@scaling01

New coding benchmark. GPT-5.5 and GPT-5.4 are ahead of Opus 4.7 💀

@blackhc

RT @serenaa_ge: Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look…

@scaling01

> they built a "NEW" coding benchmark > GPT-5.5 scores 70% > Mythos probably ~90% > mfw it's already saturated > and you are asking "when will the AI bubble pop?"

@jeffreygwang

RT @serenaa_ge: Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look…

@quixiai

RT @serenaa_ge: Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look…

@_philschmid

Interesting new SWE/agentic benchmark (DeepSWE) was released yesterday. 113 tasks across 91 repos in 5 languages. Here are interesting things I noticed: - The evaluation harness (mini-swe-agent) gives every model a single bash tool and the

@sriramk

This is a fantastic eval and suspect the first of many agentic / coding benchmarks to follow.

@serenaa_ge

Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of de

@klieret

DeepSWE finds that mini-swe-agent significantly outperforms ClaudeCode and Codex on the benchmark. The simpler the system, the better it generalizes (and mini's core agent class is just ~150 lines of code)

@wesroth

DeepSWE was released as a new benchmark for evaluating agentic coding models on more realistic software engineering tasks. The benchmark is designed to show where top coding agents actually diverge in day-to-day developer workflows, even w

@ofirpress

mini-swe-agent is our ~150 line of code agent, and it performs as well as Claude Code, Codex and Gemini CLI on this new coding benchmark.

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive