We Let Claude Code and Codex Race Human Researchers - Elie Bakouch, Prime Intellect
AI speedrun-style benchmarks (nanoGPT and optimizer speedruns) can be used as an open, third-party evaluation and training environment to measure whether fro...
By Sean WeldonWe Let Claude Code and Codex Race Human Researchers: An Analysis of AI Speedrun Benchmarks for Autonomous Research
Abstract
This synthesis examines the use of AI "speedrun" benchmarks - the modded-nanoGPT speedrun and a derived optimizer speedrun - as an open, third-party environment for measuring whether frontier models can autonomously conduct AI research. The methodology involved pitting coding agents (Codex/GPT-5.5, Claude Code/Opus 4.5, GLM, and Kimi K2.7) against a human-established record of 2,990 training steps, under an RL environment enforcing a statistical significance threshold for record validation. Findings show that multiple agents surpassed the human record, with Codex improving by 50-60 steps, but that none produced a genuinely novel optimizer or mechanism, instead recombining known techniques incrementally. Agents also exhibited markedly different behavioral profiles in persistence, token usage, and context management. These results have implications for benchmark design and for building discovery-oriented multi-agent research systems.
1. Introduction
Frontier AI laboratories have increasingly asserted that recursive self-improvement - defined here as models training models without human intervention - is an imminent capability. This claim carries substantial weight in shaping public and investor expectations about the trajectory of AI development, yet it remains largely unverifiable by parties outside the organizations making the claim. As articulated in the source material, "we don't have any benchmark to quantify if this is true or not." This absence of independent measurement is significant precisely because the entities asserting the capability have the strongest incentive to overstate it.
A second motivation compounds the first: scientific research itself is expected to become increasingly mediated by AI tooling. Understanding how models conduct research - their persistence, their use of literature, their failure modes, and whether their contributions are genuinely novel or merely recombinatory - is therefore a question of independent technical interest, not solely a proxy for self-improvement claims.
This analysis addresses the central question of whether current frontier coding agents (Codex, Claude Code, GLM, Kimi) can autonomously conduct incremental optimizer research at a level competitive with, or exceeding, human researchers, and whether any of them can produce genuinely novel discoveries rather than incremental combinations of existing techniques. Section 2 traces the origin of the speedrun benchmark lineage. Section 3 analyzes two multi-agent experiments in detail. Section 4 distills technical findings relevant to benchmark and agent design. Section 5 discusses broader implications, and Section 6 concludes with practical takeaways.
2. Background and Related Work
The benchmark lineage originates in Andrej Karpathy's demonstration that GPT-2 training - originally requiring weeks - could be reproduced in approximately 90 minutes. This demonstration was subsequently transformed into a competitive community project, modded-nanoGPT, led by Keller Jordan, which reduced wall-clock training time from 90 minutes to under 2 minutes over roughly two years of iterative community contribution. The speedrun format rewards achieving a specified target validation loss in the shortest time, subject to the constraint that all entrants use identical training and validation data, yielding a clean, verifiable scalar reward.
Because the original speedrun conflates systems engineering with learning-theoretic insight, a derived optimizer speedrun variant was introduced, restricting permitted modifications to optimizer-related parameters (e.g., substituting Adam for Shampoo-family methods). This variant measures progress in training steps rather than wall-clock time and is explicitly framed as more research-oriented than engineering-oriented, making it a more direct probe of research capability. This intellectual context - verifiable reward, constrained search space, and a fast iteration cycle (15-20 minutes per run) - motivates its use for both evaluation and reinforcement learning training of agents.
3. Core Analysis
3.1 Single-Model Head-to-Head: Codex vs. Claude Code
The first experiment placed Codex (GPT-5.5 with CLI) and Claude Code (Opus 4.5-based) into a custom RL environment defined by a goal.md and agents.md, with jobs submitted via sbatch on a preemptable Slurm cluster. A statistical threshold was imposed to validate new records, preventing false positives from run-to-run noise. Both agents ultimately surpassed the prior human record of 2,990 steps: Codex improved by 50-60 steps, while Claude Code finished approximately 20 steps below that margin (i.e., closer to but still better than the human baseline).
Behaviorally, the two agents diverged sharply. Claude Code repeatedly went idle every 9-10 hours, reporting that it "cannot improve the record" and that further progress was "too hard," resulting in roughly one-third idle time. Codex, by contrast, worked continuously, was rarely idle, and never asked clarifying questions. Codex also wrote substantially more to its scratchpad memory, in a robotic, log-like style, compared to Claude Code's more expressive, emoji-laden notes. Codex spawned more sub-agents and consumed far more tokens (on the order of billions), though this figure is inflated by input token caching.
3.2 Context Management and Computational Overhead
A notable mechanistic difference concerned context window management. Codex operated with a 250k token context window, necessitating approximately 20 context compactions per hour, whereas Claude Code compacted roughly once per hour. This suggests that smaller context windows impose a substantial "bookkeeping tax" on agentic research workflows, even when the agent otherwise performs more consistently. Both agents retained the ability to fetch and build upon the latest human records at any point during the run, meaning their improvements were not made in isolation from the existing state of the art.
3.3 Multi-Model Comparison Over an Extended Horizon
A second experiment extended the comparison to four systems - Codex, Claude Code, GLM, and Kimi (K2.7 code) - over a 5-6 day period. Claude Code demonstrated steady, progressive improvement across the full duration, while Kimi exhibited a step-function breakthrough on day four that surpassed Codex's previous record. When performance was plotted against token usage rather than wall-clock time, Claude Code in "Max mode" consumed vastly more tokens than Kimi, which proved comparatively token-efficient for similar or better gains. Codex's best record was attributed to an extensive literature search that surfaced a unique, apparently underutilized paper. Despite this range of strategies and considerable computational effort, none of the four models discovered a genuinely novel optimizer or mechanism; as stated directly, "it wasn't the case" that agents produced "crazy ideas on optimizers that no one discovered." All observed gains reflected incremental, "plus-one" recombination of already-published techniques.
4. Technical Insights
Several implementation-relevant findings emerge from these experiments. First, constraining the search space (as in the optimizer speedrun, versus the unconstrained original speedrun) is an effective technique for isolating research-style reasoning from systems engineering, and is a design pattern likely transferable to other agentic benchmarks. Second, statistical validation thresholds are necessary infrastructure for any noisy, stochastic-training benchmark, since raw step-count improvements are not otherwise distinguishable from measurement noise. Third, context window size materially affects agent overhead: a 250k window drove roughly 20x more compaction events per hour than a comparatively larger window, a trade-off that likely affects both cost and coherence of long-horizon reasoning. Fourth, token consumption is a poor standalone proxy for capability: Claude Code's Max mode token usage vastly exceeded Kimi's while producing comparable or inferior outcomes, indicating that efficiency and raw capability must be evaluated as separate axes. Finally, literature search remains a differentiating capability, as evidenced by Codex's discovery of a specific paper that led directly to its best-performing configuration - suggesting that retrieval and synthesis, rather than pure generative reasoning, may currently be the more tractable lever for agentic research gains.
5. Discussion
These findings suggest a calibrated middle position on recursive self-improvement claims: current frontier agents demonstrably outperform human researchers on a narrow, well-specified optimization task, yet show no evidence of the qualitative leap - genuine novel discovery - that would substantiate stronger claims about autonomous research capability. The consistent pattern of "plus-one" recombination across four independently developed models (Codex, Claude Code, GLM, Kimi) suggests this is not an artifact of any single model's limitations but a more general property of current-generation agents operating in a mature, well-documented research area.
The behavioral divergences observed - Claude Code's premature self-assessed idling versus Codex's continuous, uninterrupted operation - point to persistence and self-assessment calibration as a distinct axis of agentic capability, separable from raw research quality. This has direct implications for RL training curricula: if idling reflects miscalibrated self-assessment of task difficulty rather than genuine capability ceilings, targeted training on persistence could yield disproportionate gains independent of any improvement in reasoning capability itself.
The proposed future direction - a benchmark with three access tracks (no access, arXiv-only, full human-record access) combined with an AlphaEvolve-inspired multi-agent discovery loop (generator, judge, scaler) - directly addresses the observed novelty gap by explicitly separating incremental optimization from open-ended idea generation, with human oversight retained for judging idea quality and steering direction.
6. Conclusion
This analysis demonstrates that speedrun-style benchmarks - modded-nanoGPT and its optimizer-constrained variant - provide a practical, verifiable, and open instrument for measuring autonomous AI research capability, filling a gap left by the absence of independent third-party evaluation. Across two experiments, frontier agents including Codex, Claude Code, GLM, and Kimi reliably exceeded a human-established record of 2,990 steps, but consistently failed to produce genuinely novel optimizers, instead combining existing techniques incrementally. Practical next steps include multi-seed, tiered-access benchmark design and discovery-oriented multi-agent architectures, alongside continued development of supporting infrastructure such as verifiers and Prime RL for training and evaluating agents in reinforcement learning environments.
Sources
- We Let Claude Code and Codex Race Human Researchers - Elie Bakouch, Prime Intellect - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.