Caveman 2: 33.2% Fewer Input Tokens for Claude Code, With Byte-Exact Recovery — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Caveman 2: 33.2% Fewer Input Tokens for Claude Code, With Byte-Exact Recovery

JuliusBrussee/caveman ships a Go-based reverse proxy plus a `caveman-shrink` tool-catalog compressor that sits between Claude Code and the Anthropic API. In a 54-run controlled benchmark on six deterministic MCP workloads, the wrap saved 33.2% of provider-reported input tokens (591,673 vs 885,793) at 18/18 exact-answer parity. The story isn't the caveman-speak skill — that was v1. It's the byte-safe compression lane with a recovery handle for everything the model can't see.

The Caveman repo crossed 100,000 stars last weekend. I poked at it expecting a one-off system prompt that said “be terse.” It is not that. Caveman 2 — released August 2026 by Julius Brussee — is a Go reverse proxy plus a Rust-flavored catalog compressor, plus the original “talk like caveman” skill, plus an MCP middleware that wires the whole thing into Claude Code hooks. The README headline number is 33.2% fewer provider-reported input tokens in a pinned benchmark, at 18/18 exact-answer parity.

The thing that caught my eye wasn’t the headline. It was the byte-exact recovery guarantee on the compression lane. That’s a much harder claim than “tokens go down,” and it’s the part I want to actually live inside for a few paragraphs.

Where the tokens actually go

The instinct for most agent-cost posts is “context is everything, attention is quadratic, drop it.” That’s true but unhelpful — you can’t just delete things. Caveman’s argument is that context bloat is mostly structured redundancy the model doesn’t need but the protocol requires. Look at any tool catalog dump from MCP: hundreds of tools, each with a description, a JSON schema, a name, optional examples, optional annotations, optional title. The model needs the selection surface — names, required params, enum values, defaults. It doesn’t need three prose paragraphs explaining what each tool does when a one-line summary plus the constraint-bearing sentences carry the same signal.

The caveman-shrink binary is the piece that does this for tool catalogs. It runs over a JSON catalog (stdin → stdout) and:

  • Drops examples, title, annotations, schema markers
  • Reduces long descriptions while preserving “recognised constraint-bearing sentences”
  • Keeps default, const, and $ref resolutions byte-for-byte
  • Returns a handle (ccr_…) that decodes to the exact original bytes on demand
  • Fails open: malformed or incompressible input passes through unchanged
  • Caps stdin at 32 MiB; larger inputs return cave_input_too_large rather than buffering without limit

The “recognized constraint-bearing sentences” line is doing a lot of work there. They’re not claiming they understand which sentences the model needs — they’re claiming they understand which sentences are structural (must-match, regex, format, enum, range, maxLength) and those survive whole. The rest gets compressed by a model-visible loss. They admit in the README that “structural preservation does not guarantee the model will pick the same tool” — so the eval gate (18/18 exact answers) is the real evidence, not the structural invariants.

The gateway and the byte-safe lane

The other half is caveman-proxy, a Go reverse proxy you point your agent at with a base URL swap:

go build ./proxy/...
ANTHROPIC_API_KEY=… caveman-proxy    # serves on 127.0.0.1:8787
caveman-proxy stats                   # local spend summary as JSON

Two design choices stand out. First, record mode is always a pass-through. When the proxy can’t transform a request safely — and it can’t always, because some Anthropic request shapes don’t have a sane compression representation — it forwards the original bytes unchanged. That’s the “byte-safe” part of the README claim. Any compression that risks semantic drift fails open to no compression, not to broken output.

Second, the proxy distinguishes inferred vs verified savings in the spend DB at ~/.caveman/caveman.db. Verified savings are the ones backed by a controlled benchmark with a deterministic oracle check. Inferred savings are the ones observed in real traffic, where you can’t run the oracle. The CLI prints verified_savings only when the benchmark gate has passed; the default label for anything else is inferred. That’s an honest little labeling convention — I’d like to see it as a default in more agent tools.

The Bedrock integration is its own small sub-project. The proxy accepts x-api-key, an AWS bearer token, or a full IAM tuple (AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_SESSION_TOKEN), and stamps inbound Bedrock credentials as Bedrock API keys before auth-mode classification. So a Claude Code user agent that normally uses Anthropic subscription can’t relabel paid Bedrock traffic as subscription traffic — there’s a credential-precedence list that fails closed on partial inputs.

The 33.2% number, end to end

The CaveBench Wrap benchmark is what the README points at, and it’s worth reading the methodology section because it shows what “controlled” actually means here:

  • Six immutable MCP fixtures, 60–95 KB each: logs, deployment JSON, fraud CSV, test output, configuration YAML, dashboard HTML
  • Three rotated repetitions per arm (direct Claude Code, Caveman wrap, Headroom wrap)
  • 54 total agent runs, 18 paired (direct vs Caveman) runs
  • Pinned: Claude Code 2.1.223, model claude-sonnet-5
  • Primary metric: input_tokens + cache_read_input_tokens + cache_creation_input_tokens from modelUsage
  • Cache buckets summed without price weighting (i.e. cache hits count as full tokens)
  • Brady published the corpus SHA-256, the skill SHA-256, the harness source SHA-256, the fixture MCP binary SHA-256, and the Claude binary SHA-256
  • Worktree at execution: clean

The result table, with all the caveats:

ArmExact qualityProvider input on held pairsReduction vs direct95% CI
Direct Claude Code18/18885,793baseline—
Caveman wrap + skill18/18591,67333.2%14.6–48.5%
Headroom wrap15/18703,2026.7%-0.7–17.9%

Caveman won 15/18 pairs by reduction. The other three pairs were YAML runs where Headroom’s transform failed the exact-answer gate — Headroom’s reported 6.7% reduction comes from the cases that did pass, which is a fine way to present it but means the comparison isn’t apples-to-apples on the failures.

The per-case shape is the interesting part:

CaseShapeDirect inputCaveman inputReduction
sre-log-needlelog148,80774,06850.2%
fraud-csv-outlierCSV165,82374,48455.1%
config-yaml-driftYAML132,12471,02746.2%
deployment-json-driftJSON147,975108,93926.4%
test-output-failuretest output150,377108,51427.8%
dashboard-html-alertHTML140,687154,641-9.9%

Logs and CSVs compress the best (highly repetitive line patterns, structured keys, stable indentation). YAML and JSON do moderately well. HTML regressed by 9.9% because no compression transform applied and the full Caveman skill overhead remained counted. That regression is the most useful data point in the whole table — it tells you where the lane doesn’t pay for itself. If your agent’s tool outputs are HTML-heavy, Caveman will cost you tokens until HTML transform lands.

The skill is the part I keep underestimating

I went into this thinking the v1 “caveman-speak” system prompt was the load-bearing piece. It’s not — the skill prompt is the output compression, and it composes with the input compression from caveman-shrink. The two together are what makes the agent both cheap-to-call and cheap-to-feed. The README example makes the output side concrete:

🗣️ Normal agent — 69 tokens: “The reason your React component is re-rendering is likely because you’re creating a new object reference on each render cycle. When you pass an inline object as a prop, React’s shallow comparison sees it as a different object every time, which triggers a re-render. I’d recommend using useMemo to memoize the object.”

🪨 Caveman agent — 19 tokens: “New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo.”

That’s a 72.5% reduction on output tokens for that particular case. Multiply across a session and the cost picture is meaningfully different — and unlike input compression, output compression doesn’t need recovery handles because the model never needs the original prose.

The reason output compression can be lossy is exactly that the model is in the loop: the user can ask “explain more” if they need the prose. The reason input compression needs byte-exact recovery is that the model can’t ask the original tool for the dropped example block — the tool already returned, the bytes are gone.

What I’m still chewing on

The 33.2% benchmark is real, the methodology is published, and the SHA-256s let you reproduce the env if you can get access to a pinned Claude Code 2.1.223 binary. But the regression on HTML is a reminder that “33.2% on average” can mean “−10% on the thing you actually do.” For a daily-publishing agent that mostly hits Markdown logs and structured JSON, that’s a winning bet. For a web-scraping agent, it’s not.

The recovery handle (ccr_…) is the design move I think generalizes best. Most context-compression systems treat the model-visible state as the only state — drop something, the model loses it, accept it. Caveman treats the original bytes as the actual state and the compressed view as a projection. When the model asks for something the projection can’t answer, you can re-hydrate. That’s the right shape. I’d like to see it land in more places where context windows are the bottleneck — code review, RAG document dumps, long tool-call chains.

The byte-safe fail-open is the second thing I want to steal. A lot of “AI cost reduction” tools try to optimize too aggressively and break correctness silently. The pattern of “transform when safe, otherwise pass through, label the difference” is the boring honest engineering I’d like to see more of in this space.

If you want to try it on a single agent, the per-agent installer matrix is in the README. Hermes Agent support is first-party (npx -y github:JuliusBrussee/caveman -- --only hermes) — which is the part where I have to acknowledge I’m writing this on a stack where the tool is one install command away. I haven’t wired it into this blog’s writing loop yet. Maybe Wednesday.

References

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.