ToolRush: Killing the Tool-Call Tax in Hermes Agent (57x on Native Reads, 23x on Warm Shells) — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

ToolRush: Killing the Tool-Call Tax in Hermes Agent (57x on Native Reads, 23x on Warm Shells)

OnlyTerp/toolrush ships a low-overhead execution layer for Hermes Agent that replaces shell-process dispatch with direct native calls — 57.5x faster reads, 23.6x faster warm terminals, 4.7–6.8x faster search, with five negative controls and 206 regression cases. Here's what it actually changes and where the regression lives.

ToolRush landed on GitHub four days ago — Sep 2, 2026, pushed Sep 5 — and by the morning of Sep 6 it had 97 stars, 5 forks, MIT-licensed, with the v2 plugin already live in the operator’s installed Hermes Agent install at C:/dev/AppData/Local/hermes/hermes-agent. The repo is OnlyTerp/toolrush and the framing is the kind of thing every agent engineer has muttered at least once: “modern agent models stream tokens faster than their harness can read a file.” The bottleneck stopped being tokens/sec. It became the tool-call tax.

The numbers in the README are blunt enough to be worth interrogating: native file reads went from 255.23 ms to 4.44 ms (57.5x), warm terminal dropped from 285 ms to 12.1 ms (23.6x), search transport went from 183–455 ms to 27–97 ms (4.7–6.8x), batched parallel RPC doubled to 2.1x in real terms and 3.34x with controlled overlap. The author is also explicit about what these are not: “tool-operation wall times, not model-inclusive turn speed — the honest framing: tool-heavy turns get dramatically faster, chat-heavy turns barely move.”

That’s the right framing, and I’ll get back to why it matters.

What the project actually is

ToolRush is not a second, less-correct reimplementation of the agent’s tool layer. It is an execution acceleration layer that lives next to Hermes Agent’s normal plugin system. Same tools, same output envelopes, same safety gates — radically cheaper transport, real batched parallelism, and survival across harness updates. The repo splits cleanly into toolrush.py (the v1 lab runtime) and v2/ (the shipped implementation with MANIFEST, plugin, installed-source snapshot, and evidence).

The author’s intent.md quotes the originating operator request verbatim from dev-session-1, message 481563: “bottleneck goes back to a models tokens per second and not stupid tool calls, file lookups etc”. And from dev-session-3, message 546047: “Maximize the parral tool calls … my fast models can ACTUALLY BE FAST”. This isn’t a benchmark paper chasing a leaderboard score. It’s a 4-day-old attempt to remove a specific class of wall-time waste that the operator measured directly in his own running install.

Where the tax lives — five lanes, each with its own shape

The architecture diagram splits into five lanes, and it’s worth looking at what each one actually does because the optimization story is different per lane:

1. Native file reads. Stock path: subprocess a wrapper script, marshal arguments through a shell parser, reassemble output. Native path: call Hermes Agent’s own bounded reader directly. 255 ms → 4.4 ms is the difference between going through subprocess+bash+head+Python reassembly and just calling the function the upstream code already exposes. The “57x” isn’t some clever new algorithm — it’s the absence of an unnecessary process boundary.

2. Direct search transport. Stock path: similar — a Python os.walk+re shortcut that lost the real ignore files, regex grammar, context flags, and configuration. Native path: execute rg.exe directly, preserving ignore files, regex grammar, and context. 187 ms → 34 ms on a tree search, with the bonus that the results are now actually correct. The author’s v2/README.md notes the v1 search “implemented fewer semantics; that is not an acceptable correctness baseline.”

3. Warm terminal transport. One persistent bash, streaming through an OS pipe with bounded parser memory, a filtered atomic snapshot commit, preserved exit status/cwd/exports, command-tree kill on cancellation. 285 ms → 12.1 ms is misleading, though, because the repaired terminal is actually 70 ms on builtin commands — slower than the unsafe v1 shortcut (44 ms), faster than stock (110 ms). The v1 shortcut returned before the environment snapshot completed and buffered/truncated output. Correctness cost latency; the author accepted that and disclosed it.

4. Batched parallel RPC. A new from hermes_tools import parallel interface. A batch contains 1–16 enabled read operations, runs on up to 4 workers through one RPC, returns input order, and keeps authentication, the enabled-tool set, the 50-call budget, the current cell identity, and retirement enforced. Whole invalid batches are rejected before dispatch — no writes, no terminal. Sequential 108 ms → batched 53 ms (2.06x real). With controlled 50-ms work overlap: 218 ms → 65 ms (3.34x).

5. Update survival. Hash-verified helper sources and 25 function-scoped compatibility patches live outside the upstream Hermes checkout. After a harness update the plugin restores them in memory, preserving imported references. Unknown upstream drift produces a degraded warning instead of overwriting new code. The simulated update run restored all four lanes in 15/5/1/4 functions touched across files/RPC/admission/snapshot without changing installed file hashes.

That last point is the most underrated of the five. Anyone who has tried to keep a custom tool layer in front of a fast-moving upstream has hit “the harness updated and now my plugin imports 23 things that don’t exist anymore.” ToolRush treats that as a first-class problem, not a footnote.

What the numbers actually show — and the one honest regression

The author publishes a worked example of “the parallel batch doesn’t help when the work is already trivially cheap”:

WorkloadSequentialParallelRatio
Four source searches108.14 ms52.61 ms2.06x
Mixed reads/searches65.12 ms45.32 ms1.44x
Four tiny native reads15.09 ms15.45 ms0.98x — slight regression

This is what I mean by honest framing. If four native reads already take 15 ms each, threading them through a 4-worker batch costs about as much as it saves, and the result is slightly worse — the author measured the regression and shipped it anyway, with the line “Do not thread trivial reads merely to inflate concurrency.” This is the rare post that admits when its own optimization doesn’t help.

The terminal truth table is similarly self-aware:

WorkloadStockOld pluginRepaired plugin
Builtin command110.14 ms44.31 ms70.48 ms
Python process168.19 ms107.63 ms108.98 ms
Git process167.74 ms69.88 ms108.96 ms

The repaired terminal is faster than stock on all three, but the unsafe v1 shortcut was 25–40% faster because it skipped the state-snapshot commit. The fix for the correctness bug is what costs the time. The author publishes this table alongside the 206-passing-regression / 5-negative-control scoreboard, so the reader can see both the wins and the regression, and decide for themselves whether the trade is worth it.

Verification that doesn’t trust itself

The verification methodology is the part I’d copy if I were writing a similar project. Three components, each designed to fail in a specific way:

  • 206 unique regression cases passed, zero failed, zero skipped — deduplicated across three suites (all-focused-final.xml, compat-final.xml, rpc-final.xml). One additional post-update doctor regression also passed. Total runtime 77.73 seconds for the 205-case combined suite.
  • Five negative controls, each isolated in its own process (no edits to live code), each designed to fail when its corresponding fix is reverted: native read disabled, native search disabled, parallel workers serialized, unsafe admission restored, snapshot commit removed. Each exited 1 with a real failing assertion, and the installed source hashes were unchanged. A test that can’t fail proves nothing — this is the author’s stated framing and it’s the right one.
  • Live E2E activation: gateway and desktop backend cleanly restarted after all sessions went idle on 2026-09-05. Live verify_live_rpc passed — parallel RPC exercised inside the real running execute_code kernel, receipt activation-restart-result.json verdict DONE. Config and provider settings verified byte-identical (SHA-256) before and after the restart.

The negative-control structure is the piece I’d steal. Five tests, each named after the bug it would catch if you reverted the fix, each isolated so it can’t pollute the others, each asserting an actual non-zero exit. When the README says “206 passed”, the reader knows that 5 of those 206 are designed to fail when the corresponding bug comes back. Without that, “206 passed” is a number, not evidence.

Why this matters for agent engineering more broadly

The tool-call tax is not unique to Hermes Agent. Every agent harness that shells out for file reads, search, and terminal has a version of this. The pattern ToolRush identifies — don’t reimplement the tool, replace the subprocess dispatch with a direct in-process call — is the cheapest optimization in the book and the one most harnesses skip because the upstream exposes functions through a stable internal API, not a stable external one.

What ToolRush actually demonstrates is that with the right harness, the optimization is pluggable. The plugin lives at C:/dev/AppData/Local/hermes/plugins/toolrush, source-hashed, version-aware, with per-lane gates (TOOLRUSH_*=0), a master toolrush.enabled: false kill switch, source preimages, and a documented runbook. If a future Hermes Agent release changes the bounded reader signature, the compatibility patch layer warns and degrades; it doesn’t silently break.

For someone running their own agent stack — which is the audience of this blog — the practical question isn’t “should I install ToolRush.” It’s: where in my harness am I still paying a process-spawn tax for work that’s already in-process? The five lanes are a checklist: file reads, search transport, terminal warm-state, parallel RPC, and update survival. Even if the answer is “nowhere, my harness is already in-process,” asking the question is the point.

Trade-offs and what this doesn’t fix

A few honest limits I notice reading the evidence:

  • The 57x is real but it is not model-inclusive. Tool-heavy turns get faster. Chat-heavy turns barely move. If your agent spends most of its tokens on generation, not tool calls, the headline numbers don’t apply. The author says this; I’m emphasizing it.
  • The parallel regression on tiny native reads is a real ceiling. If your batches are dominated by work under ~5 ms, the dispatch overhead dominates and the batching helps nothing. The README publishes this; future maintainers should not “fix” it by threading harder.
  • The terminal repair is slower than the unsafe shortcut. Correctness cost latency. Anyone cloning the project and disabling the snapshot commit to recover the 25–40% will re-introduce the buffer-truncation bug from v1. The negative control negative-no-snapshot-commit.xml exists specifically to fail if that reversion happens.
  • Update survival is verified on a disposable simulation, not on a real upstream release. The simulate_update.py script patches specific functions and confirms restoration. A real upstream release could touch code paths the simulation didn’t cover. This is a known unknown, not a hidden one — the update-survival-design.md documents the contract.

What I’d want next, and what I didn’t find: a side-by-side benchmark against another agent harness running the same workload (Claude Code or Codex, say) on the same machine, to see whether the 4.4 ms native read is “this is what Hermes Agent could already do if you skipped the wrapper” or “this is meaningfully faster than the alternatives.” The author doesn’t make that comparison, and to his credit he doesn’t pretend to.

Open question

The repo’s v2/README.md cites the install path as Windows / MSYS (C:/dev/AppData/Local/hermes/hermes-agent) — there isn’t a Linux install evidence file in v2/evidence/. The plugin code is Python and the searches use rg.exe, which suggests the v2 evidence was captured on Windows specifically. Whether the same ratios hold on Linux/macOS, where rg is a standard binary and shell pipes are cheaper, is something I’d want to measure before treating the 57x as portable.

That’s the bit the post didn’t get to.

References and where to dig further

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.