Command Palette
Search for a command to run...

Cognition Launches FrontierCode Benchmark to Test Whether AI Code Is Mergeable; Top Model Scores 13.4/100 on Diamond Set

aiai-modelingai-research-evalsai-productsai-agents-coding 13 posts · 8 accounts

Cognition introduced FrontierCode, a software-engineering benchmark designed to test whether AI-generated pull requests would actually be accepted into an open-source repository rather than merely pass unit tests. The company said more than 20 open-source developers worked with maintainers from 36 repositories, spending more than 40 hours on each task; FrontierCode includes Extended, Main and Diamond sets with 150, 100 and 50 tasks, and the top model scored 13.4 out of 100 on the Diamond set.

FrontierCode grades beyond correctness, using unit tests, rubrics and custom verifiers to assess scope, test correctness, code quality, style and adherence to repository standards. The benchmark is aimed at a gap highlighted by METR research and Cognition's critique of SWE-bench-style grading: pull requests that pass unit tests can still include unnecessary changes or weak tests and fail human review. Cognition said its quality-control process reduced misclassification errors by 81% versus SWE-Bench Pro.

From the sources (13 posts)

@cognition

Introducing FrontierCode: a coding eval that raises the bar for difficulty & quality. Each task took 40+ hrs of work by leading open-source maintainers. Models write sloppy code that works but isn’t maintainable. Our eval is first to measu

@cognition

20+ world-class open-source developers built realistic coding tasks on repos they maintain. They define what “mergeable” means in their repo. What does it take to measure mergeability? We use a mix of unit tests, rubrics and novel verifier

@cognition

FrontierCode was built in close partnership with the expert maintainers of 36 flagship open-source repositories, like @smilingnosrati, CEO & Tech Lead @CeleryOrg (29k stars), and Martin McKeaveney, CTO of @Budibase (28k stars). Maintainers

@cognition

Rigorous quality control is important, so we built an extensive QC pipeline with adversarial testing, calibration, and multi-stage review. Every task is manually reviewed by a Cognition researcher. This reduces both the false positive and

@cognition

Tasks in the dataset have a concise problem statement with large solutions that cut across multiple files. FrontierCode’s task set is more diverse than other software engineering evals, measuring ability across a wide range of languages an

@cognition

FrontierCode has three task sets: Extended (150 tasks), Main (100 tasks) and Diamond (50 tasks). SOTA LLMs have significant room for improvement, with the top model earning a score of just 13.4/100 on our Diamond task set.

@swyx

RT @cognition: Introducing FrontierCode: a coding eval that raises the bar for difficulty & quality. Each task took 40+ hrs of work by lead…

@imjaredz

Introducing the new most important coding eval set "Is this code mergeable?" The main thing evals like SWE-bench miss is that they grade on code correctness, not quality. For FrontierCode we ask if the code maintainer will actually approv

@scaling01

Opus 4.8 is the best coding model out there FrontierCode by Cognition is probably the highest quality coding benchmark we have seen so far it moves beyond just using unit-testing for scoring, it also tests for regression safety, mechanica

@calebfahlgren

RT @cognition: Introducing FrontierCode: a coding eval that raises the bar for difficulty & quality. Each task took 40+ hrs of work by lead…

@scottwu46

SWE-Bench style grading has been the standard for years now - you ask the agent to solve an issue and then run its code on a pre-constructed unit test. The problem is that passing a unit test is only one part of writing production-ready co

@fabknowledge

RT @swyx: It's finally out!!! @METR_Evals found that more than half of SWEBench results is unmergeable slop. FrontierCode represents over…

@denizbirlikci

To understand why we built FrontierCode, read @METR_Evals's blog post on why "many SWE-bench-passing PRs would not be merged into main." A bit old now, but the point still stands. Agents often write more code — and more slop — than they sh

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive