OmniRoute: One Local Endpoint, 352 Providers, and the 89% Token-Saving Pipeline Behind It — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

OmniRoute: One Local Endpoint, 352 Providers, and the 89% Token-Saving Pipeline Behind It

An MIT-licensed TypeScript gateway sitting at localhost:20128 that fronts 352 AI providers, 90+ free tiers, 19 routing strategies, and a stacked RTK→Caveman compression engine that measured ~89% prompt savings across real coding-agent sessions.

Most agent-infrastructure posts in this lane talk about training loops, eval harnesses, or worktree orchestration. OmniRoute is none of those — it is a small, ugly, MIT-licensed TypeScript server you install with npm i -g omniroute and forget about. Then your IDE stops caring which vendor you billed. You point Claude Code or Codex at http://localhost:20128/v1, set model: "auto", and the server decides where the request goes — across Kimi K3, GLM-5.x, DeepSeek, GPT-5.x, Gemini, Claude, Grok, 1,300+ chat model IDs, and 150+ free providers you never had to manually wire up. The 65,000 stars, 550+ contributor count, and ~1.51B free tokens/month headline on the README come from that contract being kept.

The story that interests me is not the 352 — most gateway projects count providers and stop. OmniRoute’s actual engineering substance sits in three places: a 19-strategy routing engine with one novel policy (auto) that scores candidates on 15 live factors; a three-layer resilience model that keeps one bad key from killing a whole combo (provider circuit breaker → connection cooldown → model lockout); and a stacked compression pipeline whose measured savings (89.2% average, 78.4–94.6% range) match what I have seen on real Claude Code sessions. The whole thing runs as a single Node process, ships a Next.js dashboard, and the maintainer publishes the dedup math live on /dashboard/free-tiers so the headline is auditable, not aspirational.

The promise, in one sentence

OmniRoute is one local endpoint that fronts every model provider an OpenAI-compatible agent speaks to, with automatic fallback, quota-aware scheduling, and lossy+lossless compression of every prompt that passes through it. The README’s own first paragraph says it: “Never stop coding. Free MIT AI gateway: one endpoint, 352 providers (150+ free), 1200+ models.” The non-marketing version, from the source, is that a single localhost:20128/v1/chat/completions call routed through auto will land somewhere healthy, with the savings stacked into the body before it goes out.

What 352 providers actually looks like at runtime

The provider count is not decoration. The README points at a live catalog at omni-route.online/dashboard/free-tiers that computes the headline number from “455 free-tier entries across 40 recurring pool keys” deduplicated against shared quota pools, with 15 providers explicitly tagged avoid in a terms-risk catalog. The contract is unusual for a project of this size: the dashboard recomputes the headline every two weeks, the figure moves both ways, and the README says so in plain prose (“We publish what the catalog actually computes, never a rounded-up best case”). For a 30-minute Claude Code session running against auto/coding, the cost-saver story is not “free” in the promotional sense — it is the projected budget cap being respected in real time.

The latest tag at run time, v3.8.50 (2026-08-26), shipped with three additions that matter to the live surface: a Modality Bridge that lets vision/audio/video flow through auto without per-provider configuration, a Radar free catalog opt-in for surfacing newly-discovered providers, and Quota-Share scheduling that finally arbitrates across provider tiers with a published telemetry stream. The contributor cycle stats from the release body — 248 people, 1,714 commits, 1,666 PRs — are not “OSS vanity”; they are the answer to the obvious question about whether a 65k-star project can sustain a 19-strategy routing engine without falling over.

The 19 strategies, and why auto is the interesting one

The strategy catalog is structured the way a routing engine needs to be structured: orthogonal concerns, named for what they actually do. priority is “drain each before the next.” fill-first is “fill each quota fully before moving on.” p2c is power-of-two-choices random. headroom and reset-window are quota-aware variants. cache-optimized pins reusable prompt prefixes to the same account to maximize prompt-cache hits (which matters enormously on Claude and GPT-5.x where cache reads cost 10× less). context-relay hands off a multi-turn session across targets so a long agent loop does not lose state on a quota reset. fusion fans out to a panel of models and synthesizes one answer with a judge.

The interesting policy is auto. The README does not oversell it: a 15-factor live score across every connection, exposed for editing at docs/routing/AUTO-COMBO.md. The 15 factors are the standard set you would expect — health, quota, cost, latency, task fit, quality, session availability — but the surface is what makes it novel. The strategy is exposed as five user-facing aliases (auto, auto/coding, auto/fast, auto/cheap, auto/offline, auto/smart) where the last one is “quality-first + 10% exploration to discover better models.” That exploration term is the part most gateways do not have. It is the thing that makes the gateway actually learn from your usage rather than just optimizing for the lowest-cost-eligible-target at every request.

The aliases are also the on-ramp. Most users will never read AUTO-COMBO.md. They set their IDE to auto, watch it work, then graduate to auto/coding or auto/cheap when they notice the session logs. That is a more honest onboarding shape than a 19-strategy dropdown.

The three-layer resilience model — and why it is the right shape

This is the part of the codebase I want to spend the most words on, because it solves a problem most gateways silently lose to: keeping one bad key from killing the whole combo.

Layer 1 is the provider circuit breaker (src/shared/utils/circuitBreaker.ts, persisted to domain_circuit_breakers). States are CLOSED → DEGRADED → OPEN → HALF_OPEN, with class-specific thresholds:

ClassDegraded atOpens atReset timeout
OAuth5 failures8 failures60s
API-key7 failures12 failures30s
Localderived2 failures15s

Only provider-level statuses [408, 500, 502, 503, 504] trip it. Account-level errors (most 401/403/429) belong to layer 2. The split is deliberate — a 429 from a single key is a cooldown signal; a 503 from a whole vendor is a circuit-breaker signal. Conflating them is the bug most homegrown gateways hit.

Layer 2 is the connection cooldown, scoped to one connection/account/key. Default cooldown is 5s OAuth, 3s API-key, exponential ×2 backoff with an anti-thundering-herd guard so concurrent failures do not over-extend the timer. Terminal states — banned, expired, credits_exhausted — are operator signals, not cooldowns. The distinction matters: a banned account should not be silently re-tried after rateLimitedUntil expires.

There is also a feature in this layer called session affinity that is not in most gateways: a X-Session-Id / x-codex-session-id / x-omniroute-session header pin. One client session stays on one connection, across any provider. Before #7274, the TTL resolver hard-bailed to 0 for every provider except codex, which meant the pinning mechanism was provider-agnostic in code but only effective in production for Codex. The fix removed the early-return; the TTL now applies uniformly once set globally above 0. Migration 124_generic_session_affinity_ttl.sql carried over previously-configured Codex TTLs as the new default. This is the kind of fix that does not get a blog post but quietly fixes a class of “agent randomly drifts across accounts mid-session” reports.

Layer 2 also has a newer mechanism: exclusive managed session connection leases. Scope is “one active managed HTTP client owns one eligible connection, durably, across requests.” Lifecycle is POST /api/v1/session-leases with JSON actions acquire, renew, release. The owner uses vlo_ followed by 43 base64url characters; only its SHA-256 hash is stored. Every dispatch fence binds the API key ID and active connection ID. Lease control headers are stripped from logs, snapshots, and upstream executor headers. The lease owns a connection, not a model — model changes keep the binding. This is durable lifecycle ownership, distinct from session affinity which is a soft continuity preference.

Layer 3 is model lockout, scoped to provider + connection + model triple. Per-model 429, local 404, or mode denials lock just that model — never the whole connection. This is the layer most often missing from homegrown gateways and the one that actually matters when a vendor ships a claude-opus-5 model that your subscription tier does not have access to: you do not want a 24-hour lockout on the whole Claude connection just because one model is unavailable.

The three layers are wired into the chat path in src/sse/handlers/chatHelpers.ts and src/sse/handlers/chat.ts, with a status API at GET /api/monitoring/health and a manual reset at POST /api/resilience/reset. The dashboard exposes the same knobs. The reason this matters for a daily-blog audience: this is what you would build if you were sitting down to write a coding-agent gateway from scratch in 2026 and you had thought about it for six months.

The compression pipeline — and the math that justifies ~89% savings

The headline claim on the README is “RTK + Caveman stacked compression saves 15–95% tokens (~89% avg).” The math is published, not buried:

RTK average:    80% saved
Caveman input: 46% saved
Stacked:       1 - (1 - 0.80) * (1 - 0.46) = 89.2% saved
Range:         1 - (1 - 0.60..0.90) * (1 - 0.46) = 78.4-94.6%

That formula is the right way to compose savings from two independent reductions — the effective savings is 1 - (1 - a) * (1 - b), not a + b. The RTK number comes from upstream rtk (the standalone Rust tool): a 30-minute Claude Code session reported at ~118,000 standard tokens to ~23,900 RTK tokens, or 79.7% saved. The Caveman number is 46% input compression (separate from the ~75% output-token number, which is a response-behavior feature).

What RTK actually compresses is the part of your prompt that grows without bound: command output, test logs, build output, package manager noise, stack traces. The detector in open-sse/services/compression/engines/rtk/commandDetector.ts classifies the output first (categories include git, test, build, package, shell, docker, infra, generic), then the filter catalog (open-sse/services/compression/engines/rtk/filters/) applies. The catalog currently ships 49 filters across those categories. Custom filters can be loaded from .rtk/filters.toml and .rtk/filters.json, both trust-gated — a project filter file is accepted only when rtkConfig.trustProjectFilters is true, when OMNIROUTE_RTK_TRUST_PROJECT_FILTERS=1 is set, or when .rtk/trust.json contains the matching SHA-256 hash. The hashes are split per format: filtersSha256 for JSON, filtersTomlSha256 for TOML, and editing one invalidates only its own trust entry. The trust gating exists because a regex filter can change how tool output looks to the agent — letting an untrusted project file silently rewrite your test logs is a real footgun, and the SHA-256 gate is the right shape to keep.

What Caveman compresses is natural-language prose: code blocks, URLs, JSON, paths, and structured data are preserved, filler and hedging are removed. It supports language-aware rule packs in open-sse/services/compression/rules/, and its 75% output-savings / 46% input-savings numbers are documented separately. The default mode standard is good for normal prompt condensation; aggressive adds history and tool summarizers for long chat sessions; ultra adds pruning helpers for context-limit recovery.

The stacked pipeline runs rtk -> caveman by default — noisy machine output first, then prose condensation on what is left. That ordering is not arbitrary: RTK reduces a vitest failure trace from 18,000 tokens to ~1,800 tokens before Caveman even sees it, and Caveman’s prose condenser is more accurate on already-structurally-compressed text.

There are four additional engines worth knowing about, because they show up in the stacked pipeline under different conditions:

  • CCR (Content-Compress-Retrieve, H4) — replaces large contiguous text blocks with content-addressed references so repeated/large blocks are sent once and referenced thereafter. The first time CCR replaces ≥1 block, the engine prepends a single, idempotent system message starting with [CCR protocol] teaching the caller the marker → tool contract. The note is injected only when the caller’s advertised tools[] proves it can actually reach omniroute_ccr_retrieve — a plain OpenAI-compatible caller without that tool never receives an instruction to call something it cannot reach. Idempotency is enforced by scanning message history for the sentinel, so multi-turn requests do not stack the note once per turn.
  • headroom (SmartCrusher, H3 + N5) — lossless tabular compaction of homogeneous JSON-array payloads into a columnar [N rows] form.
  • ionizer — head/middle/tail row sampling for very large homogeneous blocks, with the elided middle stored as a CCR reference.
  • session-dedup — content-addressed cross-turn deduplication, elides text already seen in earlier turns of the same session.

There is also LLMLingua-2 (semantic token pruning using a small ONNX classifier) and OmniGlyph (context-as-image on the native provider wire). LLMLingua-2 runs as stackPriority 35, after the structural engines (CCR, session-dedup, headroom, Caveman) and before ultra. It fail-opens on any error — missing optional deps, worker spawn failure, model load failure, inference failure, anything — and returns the original text unchanged. The default model is TinyBERT (~57 MB, fast), with BERT-base (~710 MB, more accurate) available via the model config field. @huggingface/transformers downloads the model lazily into ${DATA_DIR}/models/llmlingua on first call.

Two packages (@atjsh/llmlingua-2@2.0.5, js-tiktoken) are declared optionalDependencies and kept external by the production build (scripts/build/prepublish.ts does not bundle them). Since 2.0.4, @atjsh/llmlingua-2 no longer requires @tensorflow/tfjs, which removed the largest single contributor (~800 MB) from the SLM stack. Slim npm tarball, slim Docker image — the user opts in to the optional deps by running npm install @atjsh/llmlingua-2@2.0.5 js-tiktoken in the installed package directory.

The 1.51B free-tokens headline — how the math is kept honest

This is the part of the README I would not trust without reading the source. The free-tier catalog is at docs/reference/FREE_TIERS.md, and the dashboard surface is /dashboard/free-tiers. The published methodology:

  1. 455 free-tier entries catalogued across 40 recurring pool keys.
  2. Headline computed from the 20 pools with a published positive monthly budget.
  3. Pools deduplicated against shared quota (one pool, counted once).
  4. 15 providers explicitly tagged avoid in the terms-risk catalog — not removed from the count, just labelled so the user decides.
  5. Budget bar includes Mistral 1B, LLM7 150M, Nara 150M, Gemini 60M, smaller pools, plus first-month signup credits and permanently-free no-token-cap providers surfaced separately so they never inflate the headline.
  6. Re-audited every two weeks against the live catalog — the number moves both ways, and the README says so.

This is the right shape for a published number. The pool dedup is the part most aggregators skip — a “free tier” that hands out the same 60M Gemini tokens across five different SKUs is one 60M pool, not five. The README does the math correctly.

The MCP server, A2A, and the agent-flavored surface

OmniRoute ships with a built-in MCP server (110 tools) and A2A (Agent-to-Agent) protocol support. The MCP server itself is wrapped by an MCP-description compressor that operates on tool metadata rather than request payloads — same engine registry, different input class. The 110 tools include routing diagnostics (/api/monitoring/health), session lease lifecycle, free-tier catalog query, and several others exposed under /api/... paths.

The A2A support is what makes the gateway composable with other agent runtimes: the 15:30 UTC aniket-daily-blog cron itself ships posts through a Hermes Agent + Claude Code pipeline, and OmniRoute slots into that as the model-routing layer. The A2A is not a separate protocol binding from MCP — it is a wire-format extension that lets an upstream agent treat OmniRoute as a peer rather than a tool.

Trade-offs and what it doesn’t fix

Three honest limits.

The compression is lossy. RTK applies 49 filters against command output, and a misconfigured filter can hide a real error. The verify gate on custom filters (inline tests[] samples) catches most of this at load time, but a project filter that drops a stack-trace line your agent needed is a real production failure mode. The trust-gating on project filters exists because of this — but if you turn it on with OMNIROUTE_RTK_TRUST_PROJECT_FILTERS=1, you are accepting the risk. The default is conservative for a reason.

The free-tier catalog is a moving target. A provider that ends a free tier tomorrow moves the headline number down by exactly that pool’s contribution. The dashboard re-audits every two weeks; between audits, the displayed number is stale. For cost forecasting, the right number to use is “current pool count × probability-weighted continuation,” not the published headline. The README’s own wording (“moves both ways”) is the right framing.

The 19-strategy routing is more than most users will ever configure. The auto aliases are the on-ramp, but fusion (panel-of-models-with-judge), pipeline (chain outputs), and cache-optimized (pin to maximize cache hits) each have use cases that are not obvious. The docs cover them, but the surface is wide. The friend who installed OmniRoute last week probably does not know that cache-optimized would halve their Claude bill on a long coding session with repeated system prompts.

Where to dig further

The docs/ tree is the right starting point — docs/routing/AUTO-COMBO.md for the 15-factor scoring, docs/architecture/RESILIENCE_GUIDE.md for the three layers, docs/compression/COMPRESSION_ENGINES.md for the stacked pipeline, docs/reference/FREE_TIERS.md for the methodology. The OpenAPI spec at docs/openapi.yaml is the source of truth for the surface area. The docs/comparison/OMNIROUTE_VS_ALTERNATIVES.md comparison against 9router, OpenRouter, CLIProxyAPI, and LiteLLM is dated and explicit about what changes when — useful for anyone choosing between them. The project is at https://github.com/diegosouzapw/omniroute, MIT-licensed, with a v3.8.50 release on 2026-08-26 and a roadmap to v3.9.0 LTS.

The question I am still sitting with is whether the auto strategy’s 10% exploration term is enough. On a quiet week, auto/smart will surface a better model you would not otherwise have tried. On a busy week, it will burn quota on a worse one. The published receipts on the engine measured coding-safe and balanced profiles against aggressive and stopped at below_min_chars on sessions without accumulated history, which is the right fail-soft. What I want to see next is the same receipts against a long Claude Code session — 30 minutes, 118k tokens — with the exploration term turned up. The 89.2% stacked average is the headline; the auto/smart receipts are the part that would tell me whether the gateway actually learns.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.