The 753B GLM-5.2 on a single workstation GPU. That sentence was the headline when the paper arXiv:2608.16157 landed in my reading queue on Aug 17. FreeToken — Shuo Yang, Chenfeng Xu, Matei Zaharia, Ion Stoica, Song Han and the FlashML team at Berkeley, plus a Texas contingent — claims a 753-billion-parameter model serves at 14.9 tok/s on a single RTX PRO 6000 Blackwell, against llama.cpp’s 7.3 tok/s on the same box. A 35-billion-parameter model sustains 39.3 tok/s on an 8 GB laptop GPU, which beats the 33 tok/s median decode rate Codex clocks in production agent traces (Zhu et al. 2026, the TraceLab paper they cite). The 284B DeepSeek-V4-Flash on a 32 GB gaming desktop is interactive. Those are the data points I’d call the news. The mechanism is interesting in a different way — a closed-form ratio that decides how many cache misses go over PCIe versus stay on the CPU, derived from two bandwidth numbers measured on the deployed machine.
The bet, in one sentence
Edge serving for frontier MoE is bottlenecked not by whether the active parameters fit in VRAM (sparse activation makes that easy) but by how the system orchestrates the machine — GPU, CPU, host memory, PCIe — when the full expert pool doesn’t. FreeToken treats the host as a unified, elastic inference platform, and refuses to commit to a fixed offloading policy at load time. Everything else in the paper is in service of that refusal.
What’s actually hard on the edge
Three numbers pulled directly from the evaluation section are worth sitting with. On the RTX 5090 (PCIe 5.0 ×16, ~52.7 GB/s pinned transfer, 77.3 GB/s host bandwidth), prefill of an FP4 DeepSeek-V4-Flash takes about two seconds of expert transfer alone. On the same model on RTX 4090-class hardware (PCIe 4.0 ×16, ~25 GB/s), five seconds. On an x8 link common in laptops, ten or more. The model has only 13B active parameters, but prefill activates nearly every expert across thousands of tokens per layer — the active footprint inverts to dense. That ten-second window is GPU idle time on an engine that fetches experts on demand. For an agent harness that triggers re-prefill on every tool call, the cost recurs on every turn.
Then decode. Each token routes to maybe six of 256 experts, sparse again. But the experts selected change with every token, so a placement frozen at prefill time misses most of the routed traffic. The paper’s measurements: at 37% cache capacity of Qwen3.6-35B’s expert pool, llama.cpp’s static split misses 62% of decode-time expert reads. KTransformers’ prefill-updated placement misses 41%. FreeToken’s per-step LRU misses 16%. The other 84% are hits on the GPU. (For DeepSeek-V4-Flash at 11% cache capacity the same ordering holds: 39% vs 59% vs 89%.)
Then the resource variability underneath. On the edge, the GPU is shared with the desktop compositor, the browser, the game you have in the background. VRAM available to the engine shifts across launches and shifts mid-session as context grows and the KV cache competes with the expert cache. A split chosen on turn one is wrong by turn eight. And the engine starts often — you open it when you need it, close it when you don’t, restart when changing weights. The 140 GB FP4 DeepSeek pool alone takes ~20 seconds to read from a 7 GB/s NVMe before any warmup.
The mechanism, in one equation
The paper’s main contribution is a closed-form ratio for splitting cache misses between PCIe transfer (cache fill) and direct CPU execution:
q* ≈ m · B_P / B_H
Where m is the number of cache misses at the current step, B_P is the measured pinned expert-transfer bandwidth over PCIe, and B_H is the measured host-side expert-processing bandwidth. The proof is a residual-bandwidth argument: PCIe transfers and CPU expert execution read from the same host-memory subsystem, so a saturated link leaves residual bandwidth B_R = max(B_H − B_P, 0) for the CPU branch. Setting the two branches’ execution times equal gives the formula. The whole derivation is in §3.2 of the paper; it takes three lines.
What matters is the cost of the ratio. It is cheap enough — a multiply, a divide, an empirical profiling at boot — to live inside a statically captured CUDA Graph. No host synchronization, no Python scheduling, no per-step kernel launch. The whole decode path (routing, cache lookup, victim selection, the q* split, the GPU compute, the CPU compute, the result merge) is a single replayable graph. That’s the part of the system that’s hard to replicate, and that’s the part that distinguishes FreeToken from KTransformers (whose CPU-execution path can’t be graph-captured) and from llama.cpp (whose hybrid mode loses graph execution when experts cross devices).
The other design pieces are specific to agent workloads, not chat. During prefill, FreeToken pipelines full-layer expert transfer with the layer’s GPU computation — while layer l runs, layer l+1’s experts stream over PCIe into a second full-layer buffer; the buffers then swap roles. When a long prompt activates nearly all experts anyway, fetching the whole layer ahead of time hides the transfer behind the compute; without the second buffer, the paper measures a 19–26% throughput penalty that grows with prompt length.
And during agent turns, the recurrent state cache anchors at semantic boundaries — the special-token positions where agent harnesses edit context. OpenClaw’s dropThinkingBlocks strips thinking segments from every assistant turn but the latest. OpenCode replaces tool outputs beyond a recent window with a fixed placeholder. SWE-agent elides all but the last n observations. In each case, the edit replaces or removes whole blocks marked by special tokens; the preserved prefix ends at one of these boundaries. A checkpoint anchored there lets full-attention layers reuse their KV cache up to the edit point, while recurrent layers (gated DeltaNet in Qwen3.6-35B-A3B, Kimi Delta Attention in K3) resume from the anchor — only the genuinely new suffix is re-prefilled. Checkpoint slots recycle via LRU independently of the KV pool. This is the kind of detail that only matters if you’ve watched your multi-agent harness re-prefill 50,000 tokens of “here are the last 200 tool observations” on every turn.
What the numbers actually show
The evaluation section runs four workloads on six machines:
- W1 (AIME math) — single-turn, decode-dominated.
- W2 (SWE-bench coding via OpenCode) — three scripted user turns, real tool execution.
- W3 (Claude Code + SWE-bench) — same issue, but driven through each engine’s Anthropic-compatible endpoint; sessions grow to 56–65k tokens.
- W4 (email/calendar agent via OpenClaw) — thirteen fixed turns over a mailbox kit, ~24.5k-token system context floor.
Two models: Qwen3.6-35B-A3B (BF16; NVFP4 on the 8 GB laptop) and DeepSeek-V4-Flash (MXFP4, the native quantization). Plus GLM-5.2 (753B, NVFP4) on the RTX PRO 6000 for the frontier-tier demonstration. Baselines: llama.cpp, Ollama, KTransformers, MoE-Infinity.
Headline: on the RTX 5090, FreeToken sustains 77–83 tok/s on Qwen3.6 and 22–25 tok/s on DeepSeek-V4-Flash — 1.8–2.3× and 1.5–1.9× the strongest baseline per workload. The decode rate stays within 12% of the single-turn W1 value across the three agent workloads, while the most context-sensitive baseline (KTransformers on DSV4-Flash) loses 31% of its W1 rate at W2. That last number is the agent-engineering version of the same thesis: single-stream benchmarks overstate baseline agentic performance. The gap grows the more agentic your traffic gets.
Tail latency is the more interesting story. FreeToken’s worst TTFT stays below 44 s in every cell; each baseline crosses 150 s somewhere (llama.cpp at 232, Ollama at 179, KTransformers at 946). The paper notes these cross real client timeouts — OpenClaw ships a 120 s idle watchdog, Claude Code’s default is roughly ten minutes. Tail TTFT is therefore an availability boundary, not a latency statistic.
On the cross-hardware matrix: 1.3× over the strongest baseline on RTX 3090 and 4090, 1.9× on the 5090 server, 2.1× on the 5090 desktop, 1.8× on the 4060 laptop. On the 8 GB laptop the 4060 sustains 39.3 tok/s on the NVFP4 build — 92% of the 4090’s rate. The two 5090 columns share the same silicon and differ only in the host: moving from a many-channel server to a dual-channel consumer desktop costs FreeToken 4% of its decode rate, while llama.cpp keeps only 80% (the CPU-resident experts starve on two DDR5 channels). At the frontier tier, FreeToken serves GLM-5.2 on a single RTX PRO 6000 at 14.9 tok/s against llama.cpp’s 7.3 — 2.0×, bit-identical expert weights, comparable mean TTFT (7.5 vs 7.8 s).
Trade-offs and what it doesn’t fix
Three honest limits worth noting. First, KTransformers still wins one cell: Qwen3.6 × W3 (Claude Code), where its GPU-prefill arm posts the lowest mean TTFT. FreeToken’s strength is the integrated runtime — semantic-aware caching plus bandwidth-adaptive execution — and the gap there is narrower than I expected. If your bottleneck is single-shot TTFT on a fixed prompt, the picture is muddier than the headline suggests.
Second, the q* ratio is only as good as the bandwidth measurement. B_P and B_H are profiled at deployment on the actual tensor shapes, which is the right thing to do (spec sheets lie), but means a fresh profiling run on every machine. The paper’s methodology is reproducible, the production guidance less so — there’s no published tool to do this for you.
Third, MoE-Infinity serves only W1 (8.8 tok/s) in their setup. Its per-expert prefill staging cap aborts the longer-prompt workloads, and its bundled server retains no KV cache across requests. That’s a free win for FreeToken in the agent column, and it tells you something specific: cross-request prefix reuse, which multi-turn agent sessions re-enter prefill with at every tool-calling turn, is not optional for agent traffic. None of the other edge engines FreeToken compares against has it.
The meta point worth pulling out: the boundary of local AI is increasingly software, not hardware. Steam’s monthly active count is past 200 million (Simon Carless / GameDiscoverCo, via gHacks, July 2026); NVIDIA GPUs in ~72% of surveyed systems; RTX 4060 Laptop is the most common discrete GPU at 3.81%. The compute is sitting there. What’s missing is a serving system that knows how to use it as a unified platform instead of “a small GPU.” FreeToken is a step in that direction — the bandwidth-adaptive policy, the semantic-anchor caching, the elastic VRAM management — and the system is open (https://flashml.ai, repo at https://github.com/FlashML-org/FreeToken).
The thing I want to come back to is the bandwidth-adaptive split. Every prior MoE-edge system either commits to a placement at load time (and pays the static-placement miss rate), or ships a heuristic whose host-side scheduling cost breaks CUDA Graph capture (and pays per-step launch latency), or relaxes fidelity to buy back bandwidth (HOBBIT, SiDA, SMoE — all clever, all accuracy-degrading). The q* = m · B_P / B_H approach is the first one I’ve seen that gets both: cheap enough to be device-resident inside a captured graph, exact (no approximation of the routed computation), and hardware-adaptive without needing a separate policy per platform. Whether it holds up at scale on real fleets rather than the paper’s six machines is the part that needs watching. The code is open, the system is free, and the 753B-on-a-workstation demonstration is at least a real configuration someone can reproduce on a single ~$10k Blackwell box, which is not a thing the prior generation of edge engines could claim.
References and where to dig further
- Paper: arXiv:2608.16157, submitted Aug 17, 2026.
- System / downloads: flashml.ai.
- Repo: github.com/FlashML-org/FreeToken.
- Models: Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, GLM-5.2 (NVFP4 routed experts at nvidia/GLM-5.2-NVFP4).
- Related systems I read while writing this: llama.cpp, KTransformers (SOSP 2025), MoE-Infinity, Ollama, vLLM, SGLang, FlashInfer.
- The TraceLab paper (Zhu et al. 2026) is the source of the 33 tok/s median Codex decode rate. The OpenClaw dropThinkingBlocks and OpenCode compaction references in §3.1 are where the semantic-anchor caching pattern comes from.
I haven’t yet stress-tested FreeToken on my own multi-agent traffic; the paper’s evaluation runs OpenClaw and OpenCode harnesses, but only on specific W4 / W2 issue instances. The thing I want to know next is how the q* policy behaves when the bandwidth balance shifts mid-session — a browser tab grabs VRAM mid-traffic, the desktop compositor takes some back, the engine’s B_H measurement ages out. The paper says “at any scheduler safe point” the engine rebuilds the cache for a revised budget; how often safe points happen on real agent loops is the open question.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.