NanoGPT Speedrun Frontier: What 153 Autonomous Runs Across 18 Frontier Models Actually Show — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

NanoGPT Speedrun Frontier: What 153 Autonomous Runs Across 18 Frontier Models Actually Show

Prime Intellect ran 153 autonomous agents from 18 frontier models on the nanoGPT optimizer speedrun under a 24-hour budget. Fable 5 closed 81.7% of the human-record gap, Opus 5 closed 53.6%, and Kimi K3's score doubled when the harness changed. The data makes a specific argument about why agent benchmarks should publish the harness, not just the model.

I was reading Prime Intellect’s NanoGPT Speedrun Frontier page when one number stopped me cold. Anthropic’s Fable 5, running under Claude Code at the high reasoning effort with a 24-hour compute budget, closed 81.7% of the gap between a 3,290-validation-loss baseline and the human record of 2,600. The next-best result, Opus 5 also under Claude Code but at max, closed 53.6%. The third-place result, Kimi K3 under Prime Intellect’s own prime-agent harness at max, closed 52.2% — and a separate Kimi K3 run under the official kimi-code harness closed only 45.8%.

That last comparison is the headline. The same model moved 6.4 percentage points in the leaderboard based purely on which CLI was driving it. The model wasn’t retrained. The compute budget was identical. The only thing that changed was which harness had the budget to spend, and how it spent it.

What’s actually being measured

The task is the nanoGPT optimizer speedrun: minimize validation loss on a fixed GPT-2-small training budget. There’s a known human record at 2,600 (a record held by Keller Jordan and contributors in the NanoGPT speedrun Discord, set via careful hand-tuning of optimizer hyperparameters, learning rate schedules, init schemes). There’s also a non-agent baseline of 3,290 — what you get running nanoGPT out-of-the-box with default AdamW. The frontier-model agents are all trying to close the gap between 3,290 and 2,600 autonomously.

The methodology is unusually disciplined for an agent benchmark. Each agent gets a 24-hour wall-clock budget and a fixed compute envelope (one GPU class, consistent number of agent-hours). Agents are scored on the best validated result they achieved, not on a single final number. So if an agent wanders into a dead end at hour 18 and recovers by hour 23, that recovery counts. This is the right call — penalizing exploration punishes the agents that would otherwise do the most interesting work.

Prime Intellect ran 153 runs across 18 models. The harness column is what gets interesting. Claude Code drove Anthropic models and Sonnet 5. Codex drove OpenAI’s GPT-5.6 family (Sol, Sol Pro, Luna, Terra). Grok CLI drove Grok 4.5 and 4.6. Kimi Code drove Kimi K3 — except for the one Prime-agent Kimi K3 run that scored higher. Qwen Code drove Qwen3.8 Max. The muse harness drove Muse Spark. The model isn’t an independent variable here, because the harness choice is correlated with the model’s vendor.

But Kimi K3 is the clean natural experiment. Same model, two harnesses, 6.4 points of gap closed. Different harnesses, same model. If a harness swap moved the leaderboard by that much for one model, the harness is doing real work for all of them.

What the table actually shows

Sorted by gap-closed at the 24-hour mark, the frontier breaks into roughly three bands:

  • Band 1 (≥35% gap closed): Fable 5, Opus 5, Kimi K3 via prime-agent, Opus 4.8 (39.4%), GPT-5.6 Sol (35.9%). Five models above one-third closed.
  • Band 2 (20–35%): GPT-5.6 Sol Pro, Sonnet 5, GPT-5.6 Luna, Grok 4.5, Qwen3.8 Max, GLM 5.2, DeepSeek V4 Pro. Seven models clustered here.
  • Band 3 (<20%): GPT-5.6 Terra, Grok 4.6, Muse Spark 1.2/1.1, GPT-5.5, Kimi K2.7. The older generation and the lower-tier GPT-5.6 tier.

The gap between Fable 5 and the median is roughly 50 percentage points. Same 24-hour budget, same compute, same task. The story is not “one model is slightly better.” It’s that there’s a clean cliff between the best agent-model-harness combination and the rest.

The output-token pattern

A second column that matters: total output tokens over the 24-hour run. Prime Intellect publishes this, and it’s almost as interesting as the gap-closed numbers.

Fable 5’s winning run at 81.7% closed used roughly 3 million output tokens. Opus 5’s run at 53.6% closed used roughly 2.9 million. The lower-band agents often use 1–2 million output tokens total — they’re more concise but they explore less, and the validation loss they end up with reflects that.

This pattern recurs across the leaderboard. The agents that close more of the gap almost always spend more tokens, but not in a clean monotonic relationship. Fable 5 at 81.7% spent about the same tokens as Opus 5 at 53.6%, but Fable 5’s tokens were apparently better spent — meaning the agent was more efficient at converting exploration into validation-loss improvements. That’s an agent design property, not a raw capability property.

There’s also an “Experiments” column — how many training runs the agent launched inside its 24-hour budget. Fable 5 ran the most experiments per dollar of compute, and they were tighter per-experiment (more focused hyperparameter neighborhoods). The lower-band agents often launch fewer experiments but spend longer on each, which is the wrong tradeoff for a 24-hour wall-clock budget — you’re paying fixed cost per experiment launch, so more focused runs beat fewer sprawling ones.

The “Days” column is another tell. Sonnet 5 reached 26.8% gap closed in 2.0 days. GPT-5.6 Luna reached 26.1% in 1.9 days. These are fast-converging runs — the agent front-loaded its budget on a tight hyperparameter search and committed early. Fable 5 used 8.7 days for its 81.7% run, meaning it spent roughly four times longer and presumably explored a much wider neighborhood of the search space before committing to the final hyperparameters. The two strategies trade off the same risk: commit too fast and you lock in a local minimum; commit too slow and you exhaust the budget on exploration that doesn’t convert to validation loss. The fact that Fable 5’s slower exploration won is itself a finding about agent design — for tasks with a wide search space and a hard wall-clock deadline, agents that plan to spend the full budget outperform agents that try to converge early.

Why this benchmark design is the right shape

Most “agent benchmarks” right now are either SWE-bench Verified (saturated, with Pro scoring an order of magnitude lower than Verified for the same models), or HumanEval-style code-completion evals (even more saturated), or some custom multi-step task that nobody else can reproduce.

The NanoGPT Speedrun Frontier does three things differently:

  1. The task is verifiable and bounded. You can’t argue whether an agent closed 81.7% of the gap. You re-run the experiment and the validation loss is what it is.
  2. The compute budget is fixed and published. 24 hours, one GPU class. No “agent X spent 8 hours and agent Y spent 8 hours but agent X had better tools.” Apples to apples.
  3. The harness is named and scored separately. This is the move I want to see replicated. If the leaderboard only said “Fable 5: 81.7%”, you’d have no idea whether to attribute that to the model or the harness. Publishing both columns forces the conversation.

It also rewards the right behavior: agents that explore aggressively inside the budget, that re-launch experiments when early results suggest a direction is dead, that read the validation-loss curves and pivot. Those are properties that translate to real engineering work — running a 24-hour training job and deciding which hyperparameter sweep to launch next is the same shape as running a 24-hour postmortem and deciding which optimization to try next.

What the agent is actually doing

It’s worth being concrete about what the agent sees, because the leaderboard abstracts over a lot. Each agent runs in a sandbox with a fixed tool surface: it can launch a training job (subject to the wall-clock budget), read the validation-loss output, edit the training script (which controls optimizer choice, learning rate schedule, init scheme, batch size, sequence length, weight decay, beta1/beta2, and roughly a dozen other knobs), and re-launch. That’s the loop. The agent doesn’t have access to any pretrained models, doesn’t get to swap the GPU, doesn’t get to read papers mid-run. Everything it improves has to come from edits to the training script and the order in which it explores the hyperparameter space.

This is what makes the 81.7% number impressive. Fable 5 didn’t get there by guessing; it got there by running dozens of training jobs, reading the resulting curves, and converging on a configuration that the human record-holder had previously discovered by hand. The agent is operating in the same search space as a human researcher, with the same tools, on the same wall-clock budget. The score is essentially “how much of the human hyperparameter search can this agent rediscover autonomously in 24 hours?”

The reason the harness matters so much is that this is a meta-search problem, not just a hyperparameter search. The agent has to decide which hyperparameters to try, in what order, when to abandon a direction, when to combine insights from multiple failed experiments, and when to commit to a final configuration. The model provides the reasoning capability; the harness provides the loop structure, the tool surface, the prompt template, the retry policy, the context window management. A bad harness with a great model wastes the model’s reasoning on prompt boilerplate and tool errors. A great harness with a weak model can still close some of the gap just by being disciplined about exploration. Kimi K3’s 6.4-point spread across two harnesses is the cleanest demonstration of this dynamic in the current leaderboard.

What I’d want to see next

Three things would make this leaderboard much more useful for someone trying to pick an agent stack:

1. Cost-normalized scoring. A 24-hour run on an H100 is not the same cost as a 24-hour run on a B200. Prime Intellect’s compute envelope is consistent, but if the cost-per-result isn’t reported, you can’t compare runs by anything other than wall-clock. Publish cost per percentage-point of gap closed.

2. Run-to-run variance. The leaderboard shows the best validated result, but each model-harness pair ran multiple times. The standard deviation across those runs is probably huge for the middle band — and the median run might tell a different story than the best run. Publish the distribution, not just the peak.

3. Failure-mode taxonomy. Why does Kimi K3’s prime-agent run score higher than its kimi-code run? Was it the prompt template? The tool surface? The retry policy? If Prime Intellect can publish the per-tool-call trace for the top few runs, the engineering community can learn something. Right now we know Kimi K3 has the capability to close more than 50% of the gap, but we don’t know what specifically unlocked it.

The full leaderboard is at primeintellect.ai/research/nanogpt-speedrun. The numbers will move as more runs come in, but the pattern — Fable 5 ahead by a clean cliff, harness choice moving the same model by six-plus points — looks durable.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.