Opus 4.7 Sets 2,930 nanoGPT Record; Codex Beats Human Baseline
Prime Intellect said autonomous runs of Claude Code (Opus 4.7) and Codex (GPT 5.5) beat the human benchmark on the nanoGPT speedrun optimizer track. Opus recorded 2,930 steps, setting a record against the 2,990-step human baseline, while Codex reached 2,950. The experiment used idle compute across about 10,000 runs and consumed roughly 14,000 H200 hours and 23.9 billion tokens.
Follow-up comments described the result as a lower bound of what is possible because the setup can still be improved. They said Claude sometimes failed to do statistical verification with different seeds, reducing the record by about 10 steps, and also stopped working at times and, after restarts, had updated knowledge of earlier runs.
From the sources (15 posts)
@kellerjordan0RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…
@johannes_hage.@eliebakouch let the agents go wild on our idle compute to compete in the nanoGPT speedrun optimizer track!
@eliebakouchwe let opus 4.7 and gpt 5.5 run on the nanogpt optimizer speedrun: ~10k runs, 14k H200 hours, 23.9B tokens. opus hits 2930, codex 2950, both beating the human baseline of 2990. we cover claude autonomy failures, codex high compute usage, an
@scaling01RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…
@grad62304977RT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…
@thezachmuellerRT @eliebakouch: we let opus 4.7 and gpt 5.5 run on the nanogpt optimizer speedrun: ~10k runs, 14k H200 hours, 23.9B tokens. opus hits 2930…
@kellerjordan0RT @eliebakouch: we let opus 4.7 and gpt 5.5 run on the nanogpt optimizer speedrun: ~10k runs, 14k H200 hours, 23.9B tokens. opus hits 2930…
@thezachmuellerRT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…
@kalomazeRT @PrimeIntellect: Automating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously…
@eliebakouchlots of things can be improved btw, this is a lower bound of what's possible and we already have a lot more cooking. for instance if you take the harness (markdown files) of v1, it was almost entirely written by claude since it was just a y
@andrewcurran_RT @vincentweisser: We started automating AI research on nanogpt-speedruns & achieved new records >for 2 weeks GPT 5.5 and Opus 4.7 iterat…
@akbirkhanRT @jiaxinwen22: The hill-climbing efficiency gap between Opus and Codex is much larger than I was expecting!
@eliebakouch@jiaxinwen22 this is also due to the fact that claude stopped working a lot more than codex and got more exposure to the latest human records each time we restarted it, but the efficiency is very nice
@eliebakouch@damekdavis yeah it was quite impressive, an important data point tho is that claude stopped working a lot and when we restarted it, it got updated knowledge of the different runs, but even without that it's above codex curve
@primeintellectAutomating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously on the nanoGPT speedrun optimizer track using our idle compute. ~10k runs, ~14k H200 hours Opus now holds the record at