Colibri: A Pure-C Inference Engine That Streams 744B MoEs From NVMe — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Colibri: A Pure-C Inference Engine That Streams 744B MoEs From NVMe

Colibri (justvugg/colibri, Apache 2.0, 32k★) runs frontier mixture-of-experts models — 744B to 2.8T parameters — on consumer hardware by treating VRAM, RAM, and NVMe as a single inference hierarchy. A single C file per model family, a five-step per-token path (route → union → place → overlap → learn), and a "JIT for weights" framing that lets the engine get faster the more you use it. v1.11.0 shipped Sep 13 with 11 tags in 30 days.

The headline number first: ./coli chat on GLM-5.2, a 744B MoE, streams at 5.8–6.8 tok/s decode on 6× RTX 5090 with full residency, TTFT ~13 s. On a 128 GB CPU-only desktop, the same model — same engine, same int4 container — runs at ~1.8 tok/s warm. On a 25 GB dev box, the proven floor: 0.05–0.1 tok/s cold. The engine is a few hundred KB of C, no GPU required, no BLAS, no Python at runtime. The model is 372 GB on disk. The claim — run a frontier MoE on hardware you already own, treat storage as another tier of memory — has not been benchmarked anywhere near this honestly before. Colibri (justvugg/colibri, Apache 2.0, 32,121★) is the most rigorous attempt I’ve seen at making the multitier-inference bet work end to end, and the maintainers publish every negative result alongside the wins.

The bet, in one sentence

A 744B MoE activates only ~40B parameters per token, and only ~11 GB of those change from token to token (the routed experts) — so the model doesn’t need to fit in fast memory, it needs to be placed. The dense part (attention, shared experts, embeddings — ~17B params) stays resident in RAM at int4 (~9.9 GB); the 19,456 routed experts (75 MoE layers × 256, plus the MTP head, ~19 MB each at int4) live on disk (~370 GB) and are streamed on demand. Limited fast memory changes speed; it must not quietly redefine the model.

What’s actually hard on the edge

Three measured numbers that ground the problem in real hardware specifics.

  • Disk read bandwidth is the bottleneck, not GPU compute, on most machines that aren’t running 6× 5090s. A single expert at int4 is ~19 MB, and a token routes to many of them per layer. Decode on the 25 GB dev box is 0.05–0.1 tok/s because the experts are streaming — the disk can’t keep up.
  • Routing has measurable structure. The repository’s #175 expert atlas measured 13,260 experts as a 3-D galaxy; 1,041 of them clustered into “specialists” by topic (poetry, law, Chinese, SQL). The hotter an expert is in your workload, the faster it should be reachable. The cache is workload-shaped, not capacity-shaped.
  • MTP — multi-token prediction heads — are easy to ship wrong. The maintainers measured the int4 MTP head at 0–4% acceptance on trivially predictable text (“count to fifty”), regenerated it at int8, and watched acceptance jump to 59%. The cache hits on draft tokens hit other experts, and the amplification (+39% experts/token, hit-rate 36→24%) made MTP a net loss on CPU streaming-MoE — both cold and warm, at 59% acceptance. The engine disables MTP by default on disk-bound configs. This is what honest measurement looks like.

The mechanism, in one equation

Every layer of every token walks the same five steps:

route → union → place → overlap → learn
  • Route: the GLM-5.2 router picks the top-K experts per token (typically 8 of 256).
  • Union: batch-union — read each unique expert once per layer, not once per token-in-batch. A layer’s routed batch-union is sent as one persistent TCP request in cluster mode, so a token doesn’t incur one round trip per expert.
  • Place: assign each expert to VRAM, RAM, or NVMe. The decision is deterministic — the same token routes to the same drive, every time.
  • Overlap: async I/O pool (PIPE=1, default) loads missing experts while resident experts compute; router-lookahead (PILOT=1) prefetches the next layer’s experts — routing is measurably 71.6% predictable one layer ahead.
  • Learn: .coli_usage records which experts your workload routes to (updated every turn) and pins the hottest ones automatically. Colibri gets faster the more you use it. The learning is per-workload, per-user, not global.

The hot-path cost that matters: each expert’s three matrices are stored adjacent and read in one pread. O_DIRECT is drive-dependent — +34% decode on Blackwell/Windows, 4.25→9.69 GB/s in iobench on a GB10, but neutral-to-negative on QLC/DRAM-less or virtualised disks. The maintainers say this out loud, not as a footnote.

// From c/colibri.c — the whole engine is one C file per model family.
// Sibling engines exist for deepseek-v4, ink, kimi-k3, qwen3, olmoe.
$ wc -l c/colibri.c
   48712 colibri.c

The 48,712-line single C file is the design choice, not a bug. The launcher picks the right binary from the model’s config.json, so the front end stays one CLI (coli chat / coli serve / coli web) across nine model families — GLM-5.2/5.3 (744B), GLM-5.3-Flash (321B, vision), Inkling (975B), Kimi K3 (2.8T), DeepSeek V4 Flash (284B), DeepSeek V4.1 Flash (552B, vision), Qwen3.8-Flash-Next (125B + 51B n-gram), Qwen3.6 (35B-A3B), and OLMoE (7B).

What the numbers actually show

Same engine, same int4 container — the hardware only changes where the experts live:

HardwareDecode tok/sTTFTNotes
6× RTX 5090, full residency5.8–6.8~13 sCUDA_EXPERT_GB=auto PIN_GB=all — disk is out of the decode path
128 GB CPU-only desktop~1.8 warm—Streaming from RAM-tier cache
Single RTX 5070 Ti laptop1.07—GPU-resident pipeline, 16 GB VRAM
25 GB dev box0.05–0.1 cold—The proven floor — disk-bound

The dual-SSD mode is the operational lever for the 25 GB case:

COLI_MODEL=/fast/glm52_i4 COLI_MODEL_MIRROR=/second/glm52_i4 ./coli chat
COLI_DISK_WEIGHTS=9,3 ...   # primary, mirror bandwidth ratio (else measured at startup)

Each expert is routed to one drive by a deterministic hash, weighted by the two drives’ measured bandwidth. A 9 GB/s + 3 GB/s pair reads experts ~33% faster than the fast drive alone. The mirror is validated at startup (per-file size + safetensors header byte-identical to primary), never written (.coli_usage, .coli_kv and sidecars stay on primary), and falls back to primary on read error (one warning, no crash). Routing never changes tokens — both copies are byte-identical. The per-run MIRROR: stats line shows GB served per drive.

The cluster mode is the horizontal extension. The coordinator keeps token generation, routing, and KV state local while disk-backed expert workers execute routed FFNs on other Macs. A layer’s routed batch-union is sent as one persistent TCP request.

./coli cluster coordinator --host 0.0.0.0 --port 8765
./coli cluster worker --model /nvme/glm52_i4 --port 9100 \
  --coordinator http://COORD:8765 --advertise-host WORKER_IP
./coli serve --model /nvme/glm52_i4 \
  --cluster-coordinator http://127.0.0.1:8765 \
  --cluster-workers HOST:PORT,...

The transport is disabled unless workers are configured, so the single-machine path stays unchanged.

The MLA KV state deserves its own note: 57× smaller than naive — 576 floats/token instead of 32,768 — and persisted across restarts (.coli_kv) so conversations reopen warm with zero re-prefill, byte-identical to an uninterrupted session. Forward validation against a transformers oracle (teacher-forcing typically 30-32/32) ties the optimization to correctness, not just throughput.

Trade-offs and what it doesn’t fix

Three honest limits, one per paragraph.

The 25 GB floor is not a deployment story, it’s a research story. 0.05–0.1 tok/s is correct but unusable. It’s the proven baseline — the project started there — but the gap between 0.1 tok/s and 1.8 tok/s is just the difference between cold and warm on the same hardware. The cache warms from your workload; if your workload doesn’t share routing structure across sessions, the cache doesn’t help. This is the difference between “Colibri runs the model” and “Colibri is useful for running the model.”

MTP — speculative decoding — is a net loss on CPU streaming-MoE, even at healthy acceptance. Issue #8 measured it cold: int4 head → 4% acceptance → draft disabled. Int8 head → 59% acceptance → 2.76 tok/fw → but 0.61 tok/s because the draft tokens routed to more experts (+39% experts/token) and dropped hit-rate (36→24%) — disk reads outweighed the saved forwards. Warm A/B: MTP=0 → 0.61 tok/s; MTP=1 → 0.47 tok/s. The MTP head works (acceptance is healthy at int8), it’s just that on this hardware class, decode is matmul/bandwidth-bound, not forward-dispatch-bound — speculation verifies ~3 tokens per forward, which increases matmul per forward instead of reducing it. The engine disables MTP by default on disk-bound configs. The lesson is that the right default depends on hardware class — what wins on Hopper might lose on a 5090.

The int4 quality story is model-specific, and one bad container poisoned a community fork. The README explicitly warns against older per-row int4 mirrors (mateogrgic/…, jlnsrk/…) — they measure ~9pp worse on quality and are the root cause of think-mode loops and never-terminating generations in #455. The gs64 container (mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp, ~372 GB) is the one to use. The MTP head must be int8, not int4 — ls -l <model>/out-mtp-* should show 3527131672 / 5366238584 / 1065950496 (int8) not the int4 sizes. The quality ablations live in docs/benchmarks.md and #108 / #81. If you’re shipping Colibri in production, the quantization question is the one to solve first.

Open hypotheses, experiments, and how to help

Colibri treats an optimization as a hypothesis until a controlled end-to-end A/B shows otherwise. The current open questions:

  • Routing history can place experts better than plain LRU (learned pins work on repeated workloads, can overfit prompts) — held-out, cross-session A/Bs needed.
  • Multiple SSDs can turn independent bandwidth into decode speed (weighted mirror/split routing is implemented and validated) — cold-cache one-drive vs two-drive GLM-5.2 runs on real, independent controllers still needed.
  • A hardware-aware planner can approach each machine’s best configuration automatically (RAM/VRAM budgets and several backends are detected today) — controlled parameter sweep across laptops, workstations, NUMA hosts, and multi-GPU systems needed.
  • Lossless or quality-bounded representations can reduce weight movement enough to matter — format and quantization ablations exist with correctness/quality gates — reproduce quality, bytes moved, latency, and cost per useful token together, not compression ratio alone.
  • Routing-aware speculation can pay before near-full residency — MTP works at int8 acceptance but loses on this engine — map the break-even surface across acceptance, expert hit rate, batch union, and draft depth.
  • CPU/GPU overlap can hide transfer and synchronization rather than merely move the bottleneck — CUDA and Metal wins exist, but fast CPUs and low residency can erase them — per-stage profiles and one-variable A/Bs across PCIe, unified-memory, and full-resident machines needed.

11 tags in 30 days (v1.6.0 Aug 12 → v1.11.0 Sep 13). The cadence tells you the project is live — the surface is moving, the maintainers are actively breaking things, and the next measurement can come from anyone willing to run the A/B.

What you actually need to install it

Two things: the program (a few hundred KB) and the model (372 GB).

mkdir colibri && tar xzf colibri-v1.8.0-linux-x86_64.tar.gz -C colibri && cd colibri
python3 coli info                         # engine ready ✓

git clone https://github.com/JustVugg/colibri && cd colibri/c
./setup.sh                                # builds + self-tests
pip install -e .                          # registers `coli` on PATH

The 11 releases, 3388 forks, 119 open issues, and per-family CUDA backends (backend_cuda.cu, backend_cuda_dsv4.cu, backend_cuda_ink.cu) tell you the maintainers are running the full matrix: GLM, DeepSeek V4, Inkling, Qwen, OLMoE, plus Metal and Vulkan paths. The Vulkan path is the only backend for hardware the vendor stacks no longer support (RX 580) and competitive with ROCm on RDNA4.

The non-obvious operational detail: partial mirrors are fine. A second SSD that holds only some shards still helps — the mirror planner ranks shards from the expert history it already learns. Run a few representative prompts first so .coli_usage reflects the workload, then mirror plan / mirror stage / mirror verify with --budget-gib and --reserve-gib. Staging never changes the primary model, copies through temporary files, preserves the requested free-space reserve, verifies every shard with SHA-256, never deletes an existing mirror shard, and atomically publishes a receipt only after the selected mirror is ready.

What’s actually interesting here

The architectural claim — multitier inference as a JIT for weights — is not new (llama.cpp has GGUF, vLLM has PagedAttention, SGLang has RadixAttention), but the implementation discipline is. The maintainers publish negative results on the same page as the wins, model-specific container warnings in the README, MTP-default-off based on measured evidence, and the O_DIRECT drive-dependency caveat that says “try it first; keep what your hardware rewards.” The model-format question (per-row int4 vs gs64 group-scaled) was a quality bug that became a documentation fix. The MTP-head-precision question (int4 head → 0% acceptance) became a default-precision fix. Every “we measured this and it lost” in the README is a paper someone could write, but instead it’s a GitHub issue with the A/B logs attached.

The bet is that frontier-scale inference becomes a systems problem, not a hardware problem. If you already own the hardware — even a 25 GB dev box — Colibri lets you hold the model instead of renting it. The cost of that is tok/s; the benefit is that the next useful optimization can come from anyone willing to measure it.

References and where to dig further

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.