StateM: 95.3% on Terminal-Bench 2.1 for $15 in API Spend, and What 'Harness Scaling' Actually Means — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

StateM: 95.3% on Terminal-Bench 2.1 for $15 in API Spend, and What 'Harness Scaling' Actually Means

Someone shipped a CLI state-machine runtime for agents this week that lifts GPT-5.5 from 83.1% to 92.1%, GPT-5.6 to 95.3% across 445 trials, and DeepSeek-V4 Flash from 82.7% to 88.1% — all without touching model weights. The total API bill for the DeepSeek runbook adaptation was $38. The final-score run on GPT-5.6 cost $15 against a $574.68 reference. The interesting story is not the score; it's that the author packaged 'make the orchestration layer a frontier-scale lever' into something you can install with `pip install -e .` and read end-to-end. I read the repo, the runbook spec, and the evaluation harness to figure out what's actually doing the work.

A repo showed up in github.com/henryqin1997/statem on Tuesday with a description that reads like a punchline to a year of agent-engineering Twitter: CLI runbook for agent long run. The paper behind it is arXiv:2608.15089, StateM ranked #1 on the Hugging Face daily papers list on Aug 18, 186 stars on the repo as of this morning, and the headline numbers are the kind that prompt the question “wait, what?“. GPT-5.5 xhigh goes from 83.1% to 92.1% on Terminal-Bench 2.1 by running it through the same runtime. The same runbook, no modifications, lifts GPT-5.6 Sol Ultra to 95.3% raw accuracy across 445 trials — every task solved at least once. The frozen runbook alone bumps GPT-5.6 Luna from 76.7% to 85.4%, comfortably above the 84.9% reference for GPT-5.6 Sol xhigh. And the total spend to adapt the runbook to DeepSeek-V4 Flash was $38, against a $574.68 reference for the GPT path. The pre-print went live two days ago; the code and reproducibility artifacts landed yesterday. I read it because the claim — “the orchestration layer is itself a frontier-scale lever” — is the single most important claim about agent engineering in 2026, and I wanted to know whether the implementation actually delivers on it.

The bet, in one sentence

StateM is a state-machine runtime for agent workflows, expressed as a versioned YAML runbook plus a CLI that executes it. The agent does not “remember” the goal; the runbook does. The agent does not decide what’s next; the runbook’s transition gates do. Verification is not a follow-up reflex; it is a precondition to leaving the current state. The whole thing is roughly 3,500 lines of Python — small enough to read in an afternoon, which is the point.

The author, Henry Qin, frames the bet as harness scaling: the idea that the execution system around the agent has the same frontier-grade leverage as the model weights, and that engineering effort spent on it is comparable in effect to a model upgrade. For 2026-agent-engineering-speak, where everyone is fighting over a four-point SWE-bench-Pro spread and trying to figure out which model’s tokens-per-dollar is actually cheapest, this is a meaningful reframe. The runbook is not a prompt. It is not a workflow engine. It is somewhere between the two — and the difference is what produces the numbers.

The runbook, copied

The README shows the canonical shape, which is worth reading carefully because it’s the abstraction doing the work:

prepare -> execute -> verify -> handoff
              ^          |
              +-- repair-+

Every state has four questions attached to it:

  • What should I do now?
  • Which transitions are legal?
  • What evidence is required before I move?
  • What happened earlier in this run?

The answers to those questions live in files on disk — not in the model’s chat history. The runtime persists the current node, transition history, hook results, evidence, timestamps, and a spec identity that ties the run to a specific runbook version. When the context window fills up, the agent can ask for a resume prompt or a compaction prompt and the runtime can generate it from the durable state — so the agent can theoretically pick up after a fresh context without losing procedural continuity. This is the trick that lets the harness survive long-horizon work without paying the full context tax on every turn.

What you actually write, in YAML, looks like this (paraphrased from examples/coding-agent.yaml):

states:
  prepare:
    description: "Read the task spec, build the working layout, identify files."
    transitions:
      - to: execute
        when: "spec-read and layout-validated"
    checks:
      - name: spec-read
        kind: file-exists
        path: ".statem/spec.md"
      - name: layout-validated
        kind: predicate
        run: "test -d .statem/working"
  execute:
    description: "Implement the spec changes."
    transitions:
      - to: verify
      - to: repair
        when: "build-failed"
    checks:
      - name: build-failed
        kind: predicate
        run: "make -n build && ! make build"
  verify:
    description: "Run the published acceptance criteria."
    transitions:
      - to: handoff
        when: "all-tests-green"
      - to: repair
        when: "any-test-failed"
    checks:
      - name: all-tests-green
        kind: shell
        run: "pytest -q"

The kind: predicate lines are the part that shouldn’t be glossed over. A predicate is a machine-checkable condition the runtime evaluates before allowing a transition. If it fails, the transition is blocked. The agent cannot “agree” to move on; the runtime has to actually see the check pass. This is what makes the runbook enforce a discipline that prompts usually only ask for.

Why this is more than a workflow engine

The README has a comparison table that names the gap. Prompt-only workflows remember the phase partially and don’t block invalid transitions. TODO lists have the same problem. CI pipelines block transitions but aren’t agent-editable. General workflow engines block transitions and survive context refresh but are rarely agent-editable. StateM is the first one in the table that checks all five boxes: it remembers phase, blocks invalid transitions, supports repair loops, survives context refresh, and is agent-editable from the CLI. That last property — agent-editable — is the unlock. The agent can ask the runtime to register a new task-specific check for the current state without mutating the shared runbook. That means the harness is programmable by the agent itself, which is the harness-scaling thesis made concrete: the agent is not just consuming the harness, it is extending it.

To put a finer point on it: most “agent harnesses” in 2026 are prompt sandwiches with optional tool wrappers. The model sees a system prompt, a few tool definitions, and a long chat history. Everything procedural — “what step am I on, what did I verify, what is the next legal move” — lives in the model’s context, and the model is asked to keep it all straight. The model is also asked to generate the next step, evaluate the previous step, and decide whether to continue, all in the same forward pass. This is what the StateM paper calls open-loop agent execution, and the failure mode is the boring one: the original goal fades, progress lives only in chat history, verification gets postponed, and a fresh session cannot reconstruct what happened.

StateM moves that procedural state out of the model context and into a lightweight, versioned runbook. The result is that the model is no longer asked to remember the procedure; it is asked to execute within a procedure. The distinction is the same one that distinguishes a shell from a heredoc: the shell has state, the heredoc is a stream. The agent stays a stream; the harness becomes the shell.

The Terminal-Bench 2.1 numbers, with the friction surfaced

The headline numbers, which I want to walk through carefully because the methodology is part of the result:

  • GPT-5.5 xhigh: 83.1% → 92.1% raw accuracy. Same model, same eval, same tasks, only the runtime changed.
  • GPT-5.6 Sol Ultra: 95.3% raw accuracy across 445 trials, with the same runbook — no model-specific tuning. All 89 tasks succeeded at least once.
  • GPT-5.6 Luna: 76.7% → 85.4% with the frozen runbook. The 84.9% GPT-5.6 Sol xhigh reference is beaten by Luna running through StateM with the same runbook GPT-5.5 was tuned for.
  • DeepSeek-V4 Flash: 82.7% → 88.1% with under $38 of API spend for the runbook adaptation. The author’s claim is that the runbook is model-class transferable — write it once for the strongest model, and a smaller model running through the same runbook can beat the larger model running without it.

The cost decomposition is the part that should make a cost-conscious team stop scrolling:

  • Final-score API spend on GPT-5.6: $15.
  • Reference (GPT-5.5 baseline to comparable accuracy): $574.68.
  • DeepSeek-V4 Flash runbook-adaptation total spend: $52.22; final accuracy exceeds the GPT-5.5 baseline by ~5 points.

This is the harness-scaling thesis in dollar terms. The orchestration layer, when treated as a primary engineering surface, is worth more than a model-tier upgrade. Whether it’s worth the engineering effort to build it is the other half of the question, and the answer on the evidence is: at a single-author, ~3,500-LoC level, yes. The author is one person. The numbers are reproducible artifacts, not a slide deck.

The thing I had to read carefully

There is one detail in the paper that I had to stare at for a while, because it determines whether the claim is “we built a better harness” or “we built a harness that’s actually different in kind.” The harness never reads the task spec the way a prompt does. It reads the task spec, parses it into a directed transition graph, and then asks the agent to execute within the graph. The graph is the unit of memory. The agent is the unit of execution. The two are decoupled by the runtime, which means the agent can be re-instantiated, the context can be compacted, the model can be swapped — and the procedural continuity survives because the graph is durable.

The practical consequence is that the runbook is also the artifact a human can audit. You can read a coding-agent runbook and tell, at a glance, what the agent is allowed to do and what it has to check before doing it. This is the move that turns “trust the agent” into “trust the runbook, watch the agent.” For multi-agent orchestration in particular — where you have several agents passing work to each other — this is the missing primitive. The current generation of agent frameworks (LangGraph, CrewAI, AutoGen, the various MCP-terminated stacks) all have some version of this, but StateM is the first one I’ve seen where the orchestration surface is a separate file from the agent code, and the file is what the agent itself edits.

There’s a curious side effect the author doesn’t emphasize. Because the runbook is versioned and the runtime records which spec identity was used for each run, the artifact of “how did this run go” is itself a structured file. You can diff two runs at the runbook level and ask: did the agent extend the runbook in this run? Which checks did it add? Which transitions did it block? This is the agent-engineering equivalent of a CI log, and it’s the kind of thing that makes postmortems into preconditions instead of folklore.

What I want to flag that the paper doesn’t

A few things worth saying out loud, because they bear on whether the result will reproduce in your shop:

  1. The benchmarks are Terminal-Bench 2.1 and BusinessBench. Both are terminal-task benchmarks — the agent ends in a terminal state, the eval script checks the result. StateM is built for this shape. If your agent’s success criterion is “the user clicked the right button” or “the deploy succeeded,” the runbook primitives still apply, but you’ll need to write your own checks. The harness is generic; the check vocabulary is not.
  2. The cost numbers are API spend, not wall-clock. A more thorough version of the experiment would also report latency and throughput. The author doesn’t, and the cost-pressure argument is strongest where the runbook is reused across many tasks — which is the case for any stable agent pipeline, but you should expect that the first time you write a runbook for a new task class, the adaptation cost is real.
  3. The frozen-runbook transfer is the strongest claim. That the same runbook, written for GPT-5.5, lifts GPT-5.6 Luna to 85.4% without modification is the lever every multi-agent team should be looking at. If you can write a runbook once and amortize it across model swaps, the model-purchase decision becomes much less load-bearing than the runbook-authoring decision.
  4. The agent-editable property is the one I’d test first. The paper says the agent can register task-specific checks from the CLI. I want to read the implementation. The released repo is small enough to audit, and the answer to “is this safe” is in the section of the code that owns the check-registration API. If you’re going to put this in production, that’s the file.

What’s next to look at

The harness-scaling thesis isn’t unique to StateM this week. The same window saw Agent Lightning v1.0 land with a complementary claim — RL on top of any agent harness via a 3,500-LoC LLM endpoint proxy, explicitly disaggregated from the harness so you don’t have to rewrite your agent to train it. The original Agent Lightning disaggregated-architecture idea has reportedly been adopted by verl Uni-Agent, AReaL 2.0, slime, and Polar; the paradigm is consolidating. Separately, Zetta ζ makes the same closed-loop argument for embodied agents — runtime policy evolution during the episode, not post-hoc reflection. Three papers in one week, all arguing that the harness is a frontier-scale lever. That’s a trend, not a coincidence.

The thing I keep coming back to is the cost decomposition. $15 of API spend on the final-score run, against a $574.68 reference. If those numbers survive independent replication on a different agent workload, every multi-agent team in 2026 is going to have to explain why their per-task spend is what it is. The vibe-coded “wire up a prompt and ship it” agent stack doesn’t have a runbook-shaped surface area to optimize. The team that ships the runbook wins the cost curve, and the cost curve is the only curve that compounds.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.