Command Palette
Search for a command to run...

Opus 4.7 Sets 2,930 nanoGPT Record; Codex Beats Human Baseline

aiai-modelingai-research-evalsai-productsai-agents-coding 15 posts · 10 accounts

Prime Intellect said autonomous runs of Claude Code (Opus 4.7) and Codex (GPT 5.5) beat the human benchmark on the nanoGPT speedrun optimizer track. Opus recorded 2,930 steps, setting a record against the 2,990-step human baseline, while Codex reached 2,950. The experiment used idle compute across about 10,000 runs and consumed roughly 14,000 H200 hours and 23.9 billion tokens.

Follow-up comments described the result as a lower bound of what is possible because the setup can still be improved. They said Claude sometimes failed to do statistical verification with different seeds, reducing the record by about 10 steps, and also stopped working at times and, after restarts, had updated knowledge of earlier runs.

From the sources (15 posts)

@kellerjordan0

RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…

@johannes_hage

.@eliebakouch let the agents go wild on our idle compute to compete in the nanoGPT speedrun optimizer track!

@eliebakouch

we let opus 4.7 and gpt 5.5 run on the nanogpt optimizer speedrun: ~10k runs, 14k H200 hours, 23.9B tokens. opus hits 2930, codex 2950, both beating the human baseline of 2990. we cover claude autonomy failures, codex high compute usage, an

@scaling01

RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…

@grad62304977

RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…

@thezachmueller

RT @eliebakouch: we let opus 4.7 and gpt 5.5 run on the nanogpt optimizer speedrun: ~10k runs, 14k H200 hours, 23.9B tokens. opus hits 2930…

@kellerjordan0

RT @eliebakouch: we let opus 4.7 and gpt 5.5 run on the nanogpt optimizer speedrun: ~10k runs, 14k H200 hours, 23.9B tokens. opus hits 2930…

@thezachmueller

RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…

@kalomaze

RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…

@eliebakouch

lots of things can be improved btw, this is a lower bound of what's possible and we already have a lot more cooking. for instance if you take the harness (markdown files) of v1, it was almost entirely written by claude since it was just a y

@andrewcurran_

RT @vincentweisser: We started automating AI research on nanogpt-speedruns & achieved new records >for 2 weeks GPT 5.5 and Opus 4.7 iterat…

@akbirkhan

RT @jiaxinwen22: The hill-climbing efficiency gap between Opus and Codex is much larger than I was expecting!

@eliebakouch

@jiaxinwen22 this is also due to the fact that claude stopped working a lot more than codex and got more exposure to the latest human records each time we restarted it, but the efficiency is very nice

@eliebakouch

@damekdavis yeah it was quite impressive, an important data point tho is that claude stopped working a lot and when we restarted it, it got updated knowledge of the different runs, but even without that it's above codex curve

@primeintellect

Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously on the nanoGPT speedrun optimizer track using our idle compute. ~10k runs, ~14k H200 hours Opus now holds the record at

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive