SGLang v0.5.18: When Draft Models Become First-Class Artifacts — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

SGLang v0.5.18: When Draft Models Become First-Class Artifacts

SGLang v0.5.18 shipped 2026-08-21 with 710 PRs from 212 contributors, and the headline isn't a single feature — it's that DSpark, DFlash, MegaMoE, and MTP all go GA together, while inclusionAI (Ant Group) uploaded the first non-Liquid DSpark draft checkpoint (Ling-3.0-flash-dspark, 1.36B params, macro mean acceptance length 5.29 across nine benchmarks). The earlier 'download SLM, run SLM' mental model is being replaced by 'download target + drafter, run with a co-designed engine,' and today's release is the moment that shift becomes the default.

I had the v0.5.18 release notes open in one tab and the inclusionAI/Ling-3.0-flash-dspark model card in another when it clicked. Both landed on Aug 21 — the SGLang release at 19:50 UTC, the drafter upload a few hours later at 05:29 UTC on the 22nd. Same week as the Liquid AI LFM2.5-DSpark blog claiming 3.2× GPU speedup and 2.9× on-device. These three things are the same thing seen from three angles.

What’s in the box

v0.5.18 carries 710 PRs from 212 contributors (release notes). The speculative decoding section alone has twenty-three separate PRs. The ones I’d flag:

  • [#31847] support inkling dspark — Inkling (Thinking Machines) becomes the first non-Liquid model whose primary vendor ships a paired DSpark drafter through SGLang
  • [#34696][#34478] Support logprobs with DSpark / Support output logprobs with DSpark — DSpark wasn’t producing logprobs at all before this; the headline inference path is now feature-complete
  • [#33459] Support logprobs with DFlash — same for DFlash
  • [#34771] Wire DFLASH aux-hidden capture into the Qwen3.5 text-only wrapper — DFlash for Qwen3.5 was a research artifact; this is the production wire
  • [#34844] Support MegaMoE for DSpark under dp attention — MegaMoE + DSpark + DP-attention in one path, the same architecture pattern Kimi-K3 uses ([#34883] pins SiTU activation for MegaMoE)
  • [#34234][#33912] Budget the DFLASH draft KV pool from its own attention geometry / fix DCP in draft KV pool sizing — DFlash draft KV pool was being mis-budgeted when DCP (Deep Context Parallel) was enabled; now it’s sized from the draft’s own attention geometry
  • [#34189] DSV4: Fix silent KV corruption when speculative draft tokens > 4 — silent KV corruption! anyone serving DSV4 with MTP depth > 4 was getting garbage KV state
  • [#34759] DSpark EP1 decode performance regression — EP=1 (single-GPU, expert-parallel off) was regressing in DSpark; fixed
  • [#33974] unified memory: support DSPARK speculative decoding + fix two NaN root causes — page hand-out was zeroing pages that DSpark had live data in, causing NaN; CuTe int32 slot-stride was wrapping; both root-caused and fixed

The diagnostic thread connecting most of these is “DSpark was working in the happy path and silently broken in the edges.” The release is the moment SGLang stopped treating DSpark as an experimental flag and started treating it as a real inference path.

The acceptance-length table that made me read the rest

The Ling-3.0-flash-dspark README has the cleanest spec table I’ve seen from a speculative-decoding drafter. Five transformer layers, hidden size 2,560, BF16, total 1,363,707,905 draft parameters, MHA with 32 query heads and 32 KV heads, target auxiliary feature layers at depths 1, 11, 23, 29, 35, vanilla Markov head with rank 256, DSpark block size 8 (verify width 9 including the bonus token). Trained with SpecForge, meant to be served with SGLang. Fine, that’s the model card.

The acceptance-length table is the part I read twice:

WorkloadAcceptance length
GSM8K6.40
MATH-5006.29
AIME 20255.56
HumanEval6.57
MBPP6.34
LiveCodeBench5.33
MT-Bench3.92
Alpaca3.51
Arena-Hard-v23.72

Macro mean across the nine workload means: 5.29. That’s the average number of tokens accepted per speculative verification step, including the target bonus token. Block size is 8, so a 5.29 acceptance length means the drafter is being rejected on average around 2.7 tokens into its 8-token proposal and the verification step costs roughly what 5.3 vanilla decodes would — except the verification step runs the target model once over all of them in parallel.

For a single-sequence AIME-2025 style reasoning trace at maybe 2000 tokens, this is the difference between 380 target-model passes and 380 target-model passes where each pass costs ≈5.3× less per token. The actual wall-clock savings depend on how much the verification pass is amortized; LMSYS measured 4.3× on HumanEval with Qwen3.5 397B-A17B DFlash block-16.

Why the drafter is published separately

There’s a structural reason inclusionAI published Ling-3.0-flash-dspark rather than baking the drafter weights into the target model. The drafter sees different inputs (auxiliary features + the target’s hidden states at the marked layers) and runs at different KV geometry. Serving them as two model artifacts under one sglang serve invocation is a co-designed path:

sglang serve \
  --trust-remote-code \
  --model-path <LING3_MODEL_PATH> \
  --tp-size <TP_SIZE> \
  --speculative-algorithm DSPARK \
  --speculative-draft-model-path <LING3_DSPARK_MODEL_PATH> \
  ...

The “draft KV pool” terminology in [PR #34234] gives away the integration shape: the draft has its own KV cache, budgeted from its own attention geometry (block size 8, 32 heads, 5 layers), separate from the target’s cache. Two pools, one engine, parallel verification. The drafter is a peer model, not a sub-component.

This is the architectural choice the LMSYS team flagged in the DFlash blog from June: “DFlash achieves higher throughput than both the baseline model and native MTP speculation in all the settings we benchmarked.” The pattern, regardless of which drafter algorithm you use, is that a small peer model running on its own KV geometry gets you better verified-token throughput than self-speculation (MTP) inside the target’s own forward pass.

The KV-corruption class of bugs

The release notes have an unusually large number of speculative-decoding fixes that turn out to be silent-state-corruption bugs. [PR #34189] is the cleanest example: “DSV4: Fix silent KV corruption when speculative draft tokens > 4.” The drafting width mattered. If you speculated more than 4 tokens ahead in DSV4 before this release, the target model’s KV cache was being silently overwritten with garbage values that didn’t match the actual tokens accepted. The accepted output looked fine because tokens get verified before commit — but anything that came after in the conversation, anything that re-used the corrupted KV pages, was operating on a poisoned cache.

This is the class of bug that looks like a benchmark regression until you trace it back. The reason speculative-decoding bugs are so dangerous is that they don’t fail loudly: they succeed in producing the right tokens, and they quietly corrupt the state machine that the next request will inherit.

Unified memory got similar treatment: page hand-out was zeroing pages that DSpark’s draft KV pool had live data in, and CuTe int32 slot-stride was wrapping under load. Both got root-caused and fixed in [PR #33974]. The fact that these three fixes all show up in one release is the signal: the speculative decoding stack reached the point where it was being used in production load patterns where these bugs actually triggered, and the production load patterns were the only place anyone noticed.

MegaMoE + DSpark + DP-attention in one path

[PR #34844] is named innocuously. What it does is support DSpark speculative decoding on MegaMoE (the Kimi-K3 / Stable Latent MoE architecture from [#34883], with explicit SiTU activation) under DP-attention (data-parallel attention, used to shard the KV heads across GPUs). This is the same combination I’m pretty sure everyone who runs Kimi-K3 at scale has been asking for, and it’s now in mainline SGLang.

The cost model of speculative decoding on MoE is the part nobody talks about. With dense models, the drafter skips the target’s largest compute per accepted block, because the drafter is much smaller. With MoE, the drafter still skips the routing decision and the routed expert dispatch — but the routed experts themselves are where the heavy FLOPs are, and the drafter doesn’t help with those. MegaMoE + DSpark is interesting because MegaMoE’s routing is more predictable than naive top-k (the latent routing path is supposed to make sparse expert selection behave more like dense activation), which means the drafter can hypothesize the routing at deeper acceptance lengths with fewer retries. The PR exists because that hypothesis survives contact with the Kimi-K3 benchmark numbers.

What isn’t here

A few release-callouts I’d have liked to see didn’t make it. The Kimi K3 MLA gate-projection fusion into the QKV-A GEMM was landed and reverted this cycle ([#33623][#34642]) — listed under “Known Issues” as not in this release. The AMD GLM-5.2 fused shared-expert append into aiter grouped-topk was also landed and reverted ([#31323][#35105]). The parallel request lifecycle tracking from gRPC generation semantics in v0.5.17 was reverted in v0.5.18 ([#34160]). These are honest acknowledgments that the engine is still churning on its hot paths.

The startup speedup headline is real though. [PR #32017] stages checkpoint pages from storage while CUDA graphs are being captured; on Qwen3-32B on H100, this made startup 8.6–11.7% faster than serial-with-prefetch and 2.38× faster than the plain default (35.6s vs 84.8s, with --startup-weight-load-mode overlap). For an engine that runs against ephemeral inference pools, that kind of cold-start ratio matters.

There’s also a clean parallel-runtime change in [PR #32313]: the TP LMHead’s allgather + scatter becomes a single all-to-all for pure-DP dp-attention. On DeepSeek-V4-Pro B200 decode, LMHead time drops from 320µs to 169µs and TPOT improves 36.97ms → 35.67ms. [PR #30700] swaps NCCL for FlashInfer MNNVL at non-fused allreduce sites — DeepSeek-V4-Flash TP4 decode on Blackwell gains up to +6.9% at small batches. Those two are the typical “free lunch” PRs that come out of a release with this many contributor-touches.

What this changes for me

The pattern I’m tracking is: the era of “download a model, run a model” is over for any deployment where latency matters. Draft models are becoming first-class artifacts, the way tokenizers and processor configs became first-class artifacts, the way quantization configs became first-class artifacts. Yesterday you’d see huggingface.co/Qwen/Qwen3.5-397B-A17B with a single safetensors blob. Today you see huggingface.co/{Qwen3.5-397B-A17B-DFlash} published in triplicate across three vendor orgs, plus a drafter independently tuned for the model. The download surface is pairing.

The next “watch” item is whoever ships a paired target+drafter pair without SGLang in the loop. llama.cpp b10577 just merged common: fix draft-mtp with embeddings (#27400) — a long-standing bug where draft-MTP models with tied embeddings produced garbage tokens. Paired with TP: enable tensor split for LFM2/LFM2MOE from the day before, GGML is shipping roughly three engine-improving PRs per day for the speculative-decoding stack. The Liquid AI LFM2.5-DSpark GGUF exists already. The same pattern is being applied in llama.cpp territory.

If your inference pipeline doesn’t yet budget for two model downloads, two KV pools, and a co-design hook between them, this is the week to start.

One thing I haven’t been able to verify is whether the Ling-3.0-flash-dspark weights come with a redistribution license or whether they’re research-use only — the model card says license: other and the inclusionAI org profile has 2,701 followers but no public license policy I’ve found. Worth checking before you ship any of this into a production inference pool.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.