I went looking for the latest gossip on r/LocalLLaMA and ended up reading the README of a project that had been quietly trending for two weeks: Paddock, a Rust inference server from a company called Truespar. The headline on their landing page — “Up to 37x faster than llama.cpp” — is the kind of claim that, from anyone else, I would have closed the tab on immediately. Then I noticed it was running 13 of 13 measured scenarios ahead of vLLM on the same FP8 checkpoint and 10 of 13 ahead of SGLang, and that the engine has no Python in the build and no CUDA toolkit in the runtime. That combination is hard to fake and worth reading further.
What Paddock actually is
Two binaries: paddock-runner (the thing that actually serves models) and paddock (a manager that bundles the Studio web UI for exploring, downloading and comparing models). The runner is the interesting part. It is written in Rust and C++, with hand-written CUDA kernels living in a separate “kernel pack” loaded over a stable C ABI at runtime. There is no Triton, no DSL, no PyTorch, no vllm-style Python runtime. The CUDA kernel pack is built independently — you can cargo build the Rust code on a machine with no GPU and no CUDA toolkit; the kernels are loaded at process start.
The first thing that struck me about the README was the explicit enumeration of what’s not there. No Vulkan. No Metal. No ROCm. CUDA only. Ampere (sm_86) and Ada Lovelace (sm_89) and Blackwell (sm_100, sm_120) only — Hopper has kernels in the tree but no measured board and no place in the release pack, so the engine refuses it unless you set PADDOCK_UNVALIDATED_ARCH=1. That kind of self-imposed constraint is unusual in this space; vLLM, SGLang and llama.cpp all want to run on everything.
The second thing was the quantification that the team chose to lead with. They don’t say “fastest inference engine in the world.” They say “13 of 13 scenarios faster than vLLM on Qwen3.8 27B FP8 at matched weight class, 1.02x to 1.19x,” with the GPU pinned (RTX PRO 6000), the dates pinned (2026-08-22/23) and the prompt-shape quadrant pinned to the industry cells. Same for SGLang: 10 of 13 ahead, +9.4% at the top, behind on 2. And against llama.cpp: 13 of 13 ahead, 1.5x to 37.5x, where the 37.5x is on a small-batch / long-output cell where llama.cpp’s lack of continuous batching really shows.
That last number is the one to remember. The 1.5–37.5x range against llama.cpp is expected if you read the existing inference-engine literature — llama.cpp’s whole engineering bet is on single-request low-latency, and it collapses as soon as you have more than a few concurrent sessions. The genuine news here is that Paddock is faster than vLLM at all 13 cells, because vLLM’s PagedAttention advantage has, up until now, been the thing you cannot easily beat without writing your own CUDA.
The kernel-pack design
The split between Rust source and CUDA kernel pack is the most interesting architectural choice. Concretely:
- Build dependency:
cargo build --releaseneeds only Rust toolchain, Node (for the embedded Studio bundle), cmake + C/C++ compiler (for the Opus codec the server compiles from source). The CUDA toolkit is required only forpacks/cuda/build.sh, which is a separate step. - Runtime dependency: the user only needs the NVIDIA driver. No CUDA toolkit on the machine. The kernel pack ships as a DLL (
.dllon Windows, equivalent on Linux) with prebuilt SASS for the supported arches. - Stable C ABI: the pack is loaded at startup. The Rust side talks to it through extern “C” functions. The README describes a CUTLASS-based FP8 GEMM (
gemm/cutgemm.cu) kept in its own translation unit, gated byPD_CUTLASS_INC, that compiles to a stub otherwise.
This is the same pattern that NVIDIA’s own TensorRT-LLM uses for its engine plan, and the same pattern that LMDeploy’s TurboMind uses for its compiled C++/CUDA serving engine. The difference is that Paddock is doing it from scratch in Rust on top, while TensorRT-LLM and TurboMind have Python wrapping C++.
The win is that the runtime is one statically-linked Rust binary with a tiny dynamic dependency for the CUDA pack. The cost is that you cannot ship a single binary that runs everywhere; you need a pack per architecture. The README is honest about which architectures have measured boards (RTX 5090, RTX PRO 6000, B200) and which only have kernels (Hopper, A100).
The quantization story
Paddock supports native FP8 (E4M3), NVFP4 (NVIDIA’s block-scale FP4), MXFP4, Q8_0 and the Q4_K_XL / Q4_K_L / Q4_K_M family. Weights load from both GGUF and safetensors. Quantization is dispatched per tensor rather than per model, because real checkpoints mix types — the Gemma 4 31B Safetensors release has FP8 attention layers and BF16 MLP layers in the same checkpoint, and any engine that bakes a quantization choice into a model file can’t load it as-is.
This is the angle that made me think the project was worth writing about rather than just bookmarking. Most inference engines still treat “the model” as a single-typed artifact. The MXFP4 / NVFP4 / FP8 dispatching is the kind of engineering that only matters if you are trying to actually serve the latest open-weight checkpoints at their native types — Qwen3.8 27B, Gemma 4 31B, Laguna 2.1 118B-A8B — and the people who built it are clearly serving them.
KV cache and the agentic workload
The README spends a paragraph specifically justifying the design for “the agentic workload: several coding-agent sessions against one box, long contexts, heavily shared prefixes.” The implementation has all the modern primitives: paged KV cache, continuous batching with chunked prefill, radix prefix caching, fair scheduling, FP8 KV storage, KV offload to host RAM, KV offload to disk, MoE expert streaming from RAM for models larger than VRAM.
That last one is the most interesting. MoE expert streaming from RAM means a model that doesn’t fit on the GPU can still serve inference at reasonable speed by paging experts in and out as the routing changes — a problem that is mostly theoretical for frontier labs but very real for anyone running 118B-A8B or larger on a single 24GB or 48GB card. Paddock’s “will-it-fit estimator” that “reports honestly instead of failing at load” is exactly the kind of thing that comes from having actually tried to run these models on boxes that were too small.
The MoE expert streaming, incidentally, is one place where the Truespar team’s earlier release truespar/siftx (a Rust metadata and PDF extraction library, MIT/Apache 2.0) shows up — they are pulling metadata into the inference loop to enrich the context, in-process rather than shelling out to a subprocess. The Studio carries an embedded WASM build of truespar/scriptor (an OOXML editor/renderer) and an embedded truespar/lector PDF viewer going open-source in September 2026, so a vision model in the Studio can render a PDF and read the rendered pages natively. That last bit — embedded WASM doc renderers — feels like a small thing, but it is exactly the friction that disappears when someone who builds doc tooling also builds the inference engine.
What I can’t verify
A few honest caveats:
- The benchmark numbers are vendor-published. Truespar published them; they ran on their hardware with their methodology. The methodology is more careful than most (aiperf as the client, same prompt-shape quadrant the industry uses, one engine resident on the GPU at a time, retokenize-on-client instead of trusting the server’s reported token count), but they are still vendor numbers. Independent reproduction will land in the next few weeks as people with RTX PRO 6000s try it.
- The “up to 37.5x faster than llama.cpp” is on a single cell. Most cells are 1.5–3x; the 37.5x is on a small-batch / long-output shape where llama.cpp’s lack of continuous batching is the dominant effect. The headline is technically correct but easy to misuse.
- Concurrency ceiling on the high end. The README says “results are promising for lower concurrency but still have a bit of work in the higher concurrency tests” for Qwen 3.8 Flash Next. So the 13-of-13 number is across the cells they ran; if you push c=128 they may not be ahead.
- Single-GPU scope. Paddock currently does not support tensor parallelism. “Models that fit on a single GPU” is the design target. That rules out serving 70B-class dense checkpoints at high throughput or anything bigger than ~70B without splitting across nodes.
- Vendor lock-in to the kernel pack. Because the CUDA is hand-written and gated by SM architecture, you cannot run Paddock on AMD or Intel accelerators even if you wanted to. That is an explicit choice.
Trade-offs the README doesn’t list
Reading further, the things not in the README are also informative:
- No Triton backend. Most modern engines have an optional Triton path because Triton lets you write kernels in Python and JIT-compile them per architecture. Paddock’s “no Python in the build” rule is incompatible with that. The flip side is that hand-written CUDA can be faster than Triton-generated CUDA for specific shapes — but only if the team keeps the kernels up to date as new model architectures land.
- No tensor parallelism. vLLM, SGLang and TensorRT-LLM all support multi-GPU TP. Paddock does not. For single-GPU Blackwell this is fine. For DGX/HGX-class deployments with multiple B200s, you would need to run multiple Paddock instances behind a router.
- No speculative decoding with draft models from outside the engine. They ship speculative decoding as a built-in feature with their own multi-token prediction, but I do not see a path to plug in a separate draft model the way SGLang lets you. Whether that matters depends on whether your draft model of choice is supported.
- Architecture coverage is narrow. sm_86 (Ampere), sm_89 (Ada), sm_100 and sm_120 (Blackwell). No sm_90 (Hopper) — even though Hopper is the most-deployed data-center inference card in 2026. They say they want contributors to help close that gap. Until then, the data-center deployment story is Blackwell-only.
Why this matters for someone running a single Blackwell box
If you are running a single RTX PRO 6000 or B200 with a coding agent fleet — Claude Code, Codex CLI, Cursor, OpenCode — Paddock’s design target is exactly your workload. Long system prompts that get re-sent hundreds of times per session. Concurrent sessions sharing one GPU. The paged KV cache with radix prefix caching is the thing that gives Claude Code its long-system-prompt throughput; the continuous batching is the thing that gives a 5-session day any throughput at all.
The cost is that you are betting on a small team. Truespar publishes security disclosures, ISO certifications and a “no telemetry, models stay on your machine” stance, and they offer free GPU time to contributors. The engineering is serious; the project is young. If you are risk-averse, stay on vLLM — the loss in single-GPU Blackwell throughput is bounded and the engine has a year of operational hardening. If you can tolerate being early to a real competitor, Paddock is the most credible new inference server I have seen land since Atlas Inference Server showed up on the NVIDIA forums in March.
I want to come back to this once independent benchmarks land on r/LocalLLaMA. The 1.02–1.19x vs vLLM is the kind of margin that disappears with one bug in the kernel pack; the 1.5–37.5x vs llama.cpp is the kind of margin that survives. If both hold up, the single-GPU Blackwell inference story just got more crowded — and that is good for everyone running agents.
References and where to dig further
- truespar/paddock on GitHub — the README, kernel pack build scripts, paddock-bench harness
- Paddock benchmarks page — the 13-of-13 board with per-cell aiperf numbers
- Paddock docs: Getting Started — install, model load, Studio walkthrough
- r/LocalLLaMA: We open-sourced Paddock — the announcement thread with practitioner pushback
- aiperf on GitHub — the load generator Truespar uses, the same one every engine uses for comparable numbers
- Atlas Inference Server on NVIDIA forums — the other recent native-inference-server competitor, single GB10 GPU
- Red Hat: llama.cpp vs vLLM — the June 2026 field-notes piece on choosing between the two incumbents, useful for the trade-off framing
- SemiAnalysis InferenceX v2 — February 2026 Blackwell-vs-AMD-vs-Hopper benchmark that grounds the hardware side of the discussion
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.