I keep a mental list of things that changed the deployment story for self-hosted models: GGUF in 2023, PagedAttention in 2023, AWQ/GPTQ quantization going mainstream, then speculative decoding shifting the TTFT–throughput frontier. vllm.cpp is the first thing in 2026 that made me re-check what “the install” actually is.
mudler — the same maintainer as LocalAI — published vllm.cpp at the start of August 2026. The headline is the size: one 66 MiB binary against a 9.1 GiB vLLM install on the same GB10. That’s a 140x difference in what you have to put on a disk to serve the same model. The tagline on the README is the one a deployment engineer actually wants to see: “Same tokens as vLLM. Same throughput. 140x less to install.” The page then spends a lot of effort proving the first two claims, which I’ll come back to.
What it actually is
vllm.cpp is a from-scratch C++20 implementation of vLLM’s serving core. No Python interpreter in the process. No PyTorch. The library is a flat C ABI — include/vllm.h is versioned at VLLM_ABI_VERSION 23 and appends fields whose zero value keeps existing behavior byte-identical. You can dlopen it from C, C++, Go, or Rust. The CLI, the OpenAI-compatible server, and the library itself are all the same binary. License is Apache 2.0 and the project is explicit about being independent and unofficial — it’s measured against vLLM, not affiliated with it.
The architecture choices are interesting precisely because they don’t try to be a new engine. Continuous batching, block-paged KV cache, automatic prefix caching, and speculative decoding come from vLLM. RadixAttention, LPM cache-aware admission, and jump-forward decoding come from SGLang. GGUF loading, compute-on-the-quantized-blocks, and the “one library behind a flat C ABI” deployment story come from llama.cpp. MLX’s GEMM is borrowed where it wins on Apple Silicon. The README puts it bluntly: “the best of each engine rather than reimplementing one of them.”
The point of that shopping list is the second thing the README does — it gates everything against vLLM as a token-for-token oracle. Outputs are checked across 27 gated architectures, byte-identical at concurrency 1 through 32, not just “close enough” in aggregate. That’s the bet: this is a port, not a re-imagination.
The numbers on Qwen3.6-27B
The headline workload is Qwen3.6-27B (NVFP4) on an NVIDIA GB10, compared against vLLM in its production graphed config (so not --enforce-eager). The throughput table, at greedy decode, closed loop:
| Concurrency | vllm.cpp (tok/s) | vLLM (tok/s) | Ratio |
|---|---|---|---|
| 1 | 86.05 | 82.32 | 1.045x |
| 2 | 159.68 | 158.03 | 1.011x |
| 4 | 292.34 | 290.31 | 1.007x |
| 8 | 508.77 | 505.46 | 1.007x |
| 16 | 801.76 | 789.16 | 1.016x |
| 32 | 1095.01 | 1076.25 | 1.017x |
vllm.cpp is ahead at every concurrency. Only c1 at +4.5% is outside the 0.5% run-to-run noise band — c2 through c32 should be read as ties, not wins. The tokens are identical either way. That’s the part that matters: switching engines doesn’t change your answers, only the install.
The other numbers worth keeping:
- 66 MiB binary vs 9.1 GiB vLLM install — measured on the same GB10. About 140x less to deploy.
- 24.88 GiB peak host memory vs vLLM’s 28.18 GiB — the same Qwen3.6-27B, the same box, no Python stack behind it. Roughly 3.3 GiB you get back for KV cache.
- Cold start to first
/health: 36.5 s vs vLLM’s 221.5 s — provisional but reproducible via.agents/benchmark-record.md. That’s a 6.1x faster cold start, which matters in two places: autoscaling with low min-replicas, and edge deployments where you actually pay the cold-start tax. - MXFP4 parity — Qwen3-8B MXFP4 runs W4A16 Marlin by default and decodes 45.45 vs 41.94 tok/s, again token-for-token identical to vLLM.
CPU and Apple Silicon, with caveats
The same binary runs CUDA, CPU, Metal, and Vulkan from one source tree. ROCm and Tenstorrent are growing. Against llama.cpp on the same GGUF file (CPU):
| vllm.cpp | llama.cpp | ratio | |
|---|---|---|---|
| prefill | 223.8 tok/s | 177.3 | 1.18x |
| decode | 24.7 tok/s | 25.4 | 0.97x (tie) |
| peak memory | 2.83 GiB | 2.80 GiB | 1.01x |
Prefill is meaningfully ahead; decode is a tie. Tokens are byte-identical to llama.cpp’s greedy decode, which is the only correctness claim that matters on a quantized file. There’s a single big caveat the README flags in bold: every llama.cpp denominator here is SUPERSEDED. Those numbers came from 237ad9b96, a local-only fork the team kept “65 performance commits deep.” The pin is now stock b10451 and each figure is owed a re-take that can move it either way — see issue #1003. So treat the 1.18x prefill advantage as a direction, not a settled fact.
Against MLX-LM on an Apple M4, warm b=1:
| vllm.cpp | MLX-LM | ratio | |
|---|---|---|---|
| prefill TTFT | 524.5 ms | 532.6 | 1.5% ahead |
| decode | 27.23 tok/s | 27.85 | 97.8% |
| warm total | 24.37 tok/s | 24.96 | 97.6% |
That 2.4% gap is real, not noise: across six interleaved runs the spread was 0.12% for vllm.cpp and 0.34% for MLX-LM. All of it sits in decode — about 0.81 ms per token — and the team says they know where it goes. Under the hood, their Metal GEMM runs at 97% of MLX’s own (3.91 TFLOP/s mma issue rate), the decode GEMV streams weights at 83% of the part’s memory-bandwidth peak, and moving prefill attention onto the matrix units was worth 4.3x. Two models tested (OPT-125m, Qwen3-0.6B), 18 of 75 ops native, and the 97.6% needs the optional MLX GEMM provider shape-gated to prefill. Indicative, not binding.
What you actually get beyond vLLM parity
The list of vLLM-parity features plus the engine-specific additions is long and worth reading directly from docs/FEATURES.md, but the parts that matter for an inference engineer:
- GGUF as a first-class citizen. Same quantized files llama.cpp uses. On CPU, compute directly on the compressed blocks (Q4_0, Q8_0, Q3_K, Q4_K, Q5_K, Q6_K) with no BF16 expansion.
- SGLang knobs, opt-in. RadixAttention / prefix caching, LPM cache-aware scheduling, jump-forward decoding, and custom logits processors are documented toggles. Each defaults to today’s behavior, so a build that sets none of them is byte-identical to one built without them.
- Three speculative decoders: MTP, block-diffusion DFlash, and draft-free ngram. MTP is token-identical to vLLM’s MTP and about 4% faster at c1 on Qwen3.6-27B-NVFP4. DFlash runs about 2x over spec-off but stays below vLLM’s throughput — an open bf16 acceptance floor.
- Structured output including GBNF, 38 tool-parser families, image, video, and audio input, music generation through
POST /v1/audio/speech, external KV offload, and Prometheus metrics. All in the samedlopen-able library. - 40 registered architectures as of v0.0.2. ROCm and Tenstorrent are growing.
The bit that caught my eye in particular is the C ABI. The vLLM project has historically been a Python-first service — you don’t embed it, you run it. vllm.cpp’s include/vllm.h with VLLM_ABI_VERSION 23 is a serious bet that inference will become a library people link against the same way they link against libcurl or libpng. That’s a different shape for the deployment story than “a long-running server you put behind a reverse proxy.”
What the numbers don’t show
Three things I want to flag before this sounds too good.
One GPU, one CPU path. All the speed claims are proven on a GB10 (sm_121a) for GPU and a CPU path that matches or beats llama.cpp on GGUF. The cross-architecture work is the v0.0.2 release, and the README is honest that “most other architectures are speed-pending, and say so.” The correctness gates cover 27 architectures; the speed numbers do not. If you’re on H100, MI300X, or a 4090, the picture is probably similar but is not measured here.
Heavy development, expect breakage. The project is pre-1.0 and the README warns that internals, CLI flags, and server behavior can change between commits if you track main. The one thing they keep disciplined is the C ABI. That’s the right thing to keep stable, but it also means everything outside that header is fair game.
The competition is also moving. llama.cpp is itself working on paged KV and a continuous-batching scheduler modeled on vLLM (see ggml-org/llama.cpp discussion #21961). If llama.cpp closes the concurrency gap, a lot of the vllm.cpp story becomes “you get vLLM’s serving quality, but at the cost of giving up llama.cpp’s reach.” SGLang is also adding GGUF support and the deployment story there is going to look different in six months. jmaczan’s tiny-vllm is a teaching C++/CUDA port of vLLM that’s been around since May — it’s a different bet (educational, not production) but it’s the kind of thing that tells you the field is consolidating.
The thing I’m sitting with: the 6.1x cold-start win is bigger than the throughput numbers, and the 140x install-footprint win is bigger than the cold-start win. Most inference servers in production are not throughput-bound at concurrency 32. They’re cold-start bound at concurrency 1, or they’re constrained by how many replicas you can fit on a disk image, or by how fast a new node can join a cluster. That’s where vllm.cpp’s numbers are the most honest “yes this changes the deploy” statement I’ve read in a while.
The version I want to keep an eye on is the next one that adds a second GPU to the speed table. If the GB10 result replicates on H100 and MI300X, the vLLM project itself will probably have to answer the install-size question. Until then, treat the 140x number as a real and reproducible result on the device it was measured on, and a strong hint elsewhere.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.