Command Palette
Search for a command to run...

Opus 4.7 Sets nanoGPT Record While Codex Beats Human Baseline

aiai-modelingai-research-evalsai-productsai-agents-coding 3 posts · 3 accounts

Claude Code (Opus 4.7) and Codex (GPT 5.5) beat the 2,990-step human baseline after running autonomously on the nanoGPT speedrun optimizer track. Opus posted 2,930 steps to set the record cited in the experiment, while Codex reached 2,950.

The test used idle compute across about 10,000 runs, consuming roughly 14,000 H200 hours and 23.9 billion tokens. The accompanying write-up said it also detailed Claude autonomy failures and Codex's high compute usage.

From the sources (3 posts)

@kellerjordan0

RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…

@johannes_hage

.@eliebakouch let the agents go wild on our idle compute to compete in the nanoGPT speedrun optimizer track!

@eliebakouch

we let opus 4.7 and gpt 5.5 run on the nanogpt optimizer speedrun: ~10k runs, 14k H200 hours, 23.9B tokens. opus hits 2930, codex 2950, both beating the human baseline of 2990. we cover claude autonomy failures, codex high compute usage, an

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive