The 104 GB Qwen3.8-Flash-Next 4-bit weights landed on my machine’s free-space list before slotstream was released. I’d been telling myself the model “wasn’t for me yet” — that’s the polite phrase for “I have 32 GB of RAM and 110 GB of weight file.” Show HN landed on Sep 1 with 114 points in the first day, and a single sentence from the README reframed the constraint: keep the 3.8 GB trunk resident, stream experts through a fixed slot pool, size the pool to what the machine actually has. The storage layer, not the GPU memory or the disk bandwidth, is the rate limit.
That sentence is wrong in the obvious direction and right in a less obvious one. Yes, SSD throughput sets the floor. But the real architectural choice is the memory design: experts get loaded from SSD straight into memory the GPU already addresses, and the slot pool sizes against unified memory. A discrete GPU has two separate budgets with a bus between them — that’s a second engine with a second cache tier, not a port, which is why slotstream is Apple Silicon only.
The numbers, as measured
The 48 GB M5 Pro row is real hardware. The smaller-Mac rows come from slotstream doctor --sim-ram N:
| your Mac | slotstream takes | warm decode |
|---|---|---|
| 8 GB | 8.1 GB (the floor) | ~3 tok/s + paging warning |
| 16 GB | 10 GB | ~4 tok/s |
| 24 GB | 16 GB | ~8 tok/s |
| 32 GB | 22 GB | ~9 tok/s |
| 48 GB+ | 33 GB | ~12 tok/s |
Engine start is ~2 s — only the 3.8 GB trunk loads. Peak memory on the 48 GB run is 32 GB (auto-sized, you can cap it). Weights on disk: 105.3 GB across 25 files. The last 1.5 GB file is a draft head; without it, speculative decode is off but everything else works.
Auto-sizing never takes the whole machine. With a browser open, auto takes less and prints the plan at startup. Run slotstream doctor before downloading anything.
The mechanism
# conceptual, not the actual Swift code — read this as the cache shape
TRUNK_BYTES = 3.8 * 2**30 # stays resident, never swapped
SLOT_POOL_BYTES = machine.ram * 0.5 # sized at startup, never grows
N_SLOTS = SLOT_POOL_BYTES // EXPERT_BYTES
def forward(token):
# 1. compute which experts are needed (top-k router)
# 2. if expert not resident: evict LRU slot, mmap that range from SSD
# 3. swap time dominates only at slot-pool saturation
# 4. never touches the 3.8 GB trunk — that's the working set
The key insight: a MoE model with 125B params and ~6B active is not a 125B model for memory purposes. It’s a 3.8 GB trunk plus a pool of expert shards, of which you touch a small active subset per token. The traditional loader tries to fit all of it; the OS goes to swap; the first token takes 30+ seconds. Slotstream’s memory design treats the SSD as the primary expert storage and unified memory as the cache. The floor at 8.1 GB is what the trunk plus minimum slot pool needs to function without paging.
The repository has a MEASUREMENTS.md file where the author walks through every number that failed, starting with “M07: the naive path fails — why slotstream exists.” The first measurement on the 48 GB Mac using the stock loader showed the machine going into 48 GB of swap before the first token. That’s the counterfactual: without the memory redesign, the same hardware doesn’t run the model at all.
What this actually changes
The implication, spelled out: any 100B+ open-weight model is now “runs on a 16 GB Mac with a 512 GB SSD” territory. The streaming layer, not the weights, not the VRAM, is the engineering surface. vLLM and llama.cpp are now competing on a different axis — not just kernel speed, but how well they can treat SSD as part of the memory hierarchy.
Practically: the install is one command, the binary is a single Swift file, and it speaks Ollama and OpenAI chat APIs. Existing tools work unchanged. The download prints the size and your free disk and refuses outright if the disk can’t hold it. Releases are built by CI with signed provenance, so gh attestation verify works on the asset.
curl -fsSL https://raw.githubusercontent.com/carloslfu/slotstream/main/install.sh | sh
slotstream pull # 105 GB, asks first
slotstream run # Ollama-compatible API
What it doesn’t fix
Apple Silicon only, and not by accident. The memory design assumes unified memory — experts go from SSD into memory the GPU already addresses, and the pool sizes against the one memory the OS, your apps, and the GPU all share. A discrete card has two separate budgets with a bus between them; a Linux or Windows build would be a second engine with a second cache tier, not a port. It isn’t on the roadmap.
The smaller-Mac rows are estimates from the curve, not measurements. Smaller Macs also have slower SSDs, so the 8 GB and 16 GB rows are upper bounds in practice. The author is explicit: only the 48 GB row is real hardware. docs/HARDWARE.md is a ten-minute procedure to measure your SSD bandwidth and get your own row into the table.
The auto-sizing never takes the whole machine, but it also doesn’t know your workload. With a browser open, auto takes less and says so. With a long-running generation, the slot pool can saturate even on the 48 GB tier — the README is honest that the active subset per token is small, but during a long generation with heavy expert reuse, you’ll see more slot evictions and slower decode than the steady-state 12 tok/s.
Where I’d push this next
The model-specific code path is hard-coded for Qwen3.8-Flash-Next. A second model would require re-deriving the trunk size and the expert layout. The repository’s “Related projects” section points at llama.cpp for other platforms and other models — the right framing is that slotstream is a memory-design proof, not a universal inference engine.
The honest thing to say about a 4 tok/s decode rate is: it’s fast enough to feel responsive, slow enough to discourage anything but chat-shaped workflows. If you’re trying to run a 100B model locally for code generation that takes 30 seconds per response, you’re going to feel the SSD. For “let me ask this model a quick question about a document I have open,” 4 tok/s is fine.
I haven’t run this yet. The 105 GB download is still sitting in my pending queue behind the 30 GB one. But the HN thread’s comment section is doing what good comment sections do — pointing out edge cases the author hasn’t tested (eMCP sleep behavior, what happens when Time Machine kicks in mid-generation, how it interacts with M5 Pro’s SSD wear-leveling over weeks of use). Those answers will be in MEASUREMENTS.md next month.
What I want to know before I run it
Three things aren’t answered in the README and probably matter in practice. First: what’s the slot eviction policy during a long generation? The author’s framing — fixed pool, evict LRU when an expert isn’t resident — is correct for steady-state but underspecified during expert-heavy workloads. If the model’s router converges on a small active subset per token (which is what makes MoE fast), the pool stays warm and you get 12 tok/s. If the router diverges, you start hitting disk on every forward pass and the decode rate collapses. The MEASUREMENTS.md has the steady-state numbers; nobody’s published the worst-case yet.
Second: what happens during SSD contention? Time Machine snapshots, Spotlight indexing, Photos sync — these all hit the same SSD slotstream is streaming from. A long generation during a Time Machine window will see decode rate drop to whatever your SSD bandwidth is minus the snapshot’s. The author is aware (the install prints a free-disk check) but the README doesn’t quantify the impact.
Third: how does M5 Pro’s SSD wear-leveling interact with 105 GB of stream reads per install, plus ~GB-scale stream reads per session? The first install is fine — modern SSDs handle 100+ TBW without issue. But if you re-download every other day because of pre-release breakage, that’s a different calculation. The author’s release cycle is monthly, so in practice this isn’t a problem.
How the architecture compares to llama.cpp
llama.cpp has been doing SSD-aware inference for years via mmap and the -mlock / --no-mmap knobs. The difference is what slotstream is doing with MoE specifically: llama.cpp’s mmap path treats the model file as a flat address space and pages experts in/out on demand, but the slot pool size is not as carefully tuned to the active subset per token. Slotstream’s measured numbers — 12 tok/s warm decode on a 48 GB Mac — are higher than what I’ve seen published for llama.cpp on the same hardware and model, though the comparison isn’t apples-to-apples because llama.cpp’s MoE path has been changing weekly.
The interesting question is whether llama.cpp adopts the slot-pool pattern, or whether slotstream remains an Apple-Silicon-only special case. The author is explicit: it’s not a port. On another platform, llama.cpp is the right starting point. So in practice, the field bifurcates: Apple Silicon users get slotstream, NVIDIA/Linux users get llama.cpp, and Windows/Linux discrete-GPU users get… whatever llama.cpp does on those paths.
The proof point worth watching: when Qwen3.8-Flash-Next 4-bit lands on a 16 GB Mac at 4 tok/s and someone uses it for real work, does the design pattern spread? My bet: yes, within six months. The memory-design insight is general — any MoE model with a small active subset is now a candidate for SSD-streamed expert offload, not just Qwen3.8-Flash-Next.
References and where to dig further
- Slotstream on GitHub — single Swift binary, MLX + Metal, Ollama/OpenAI API compatible
- Show HN thread — 114 points in 24h on Sep 1, the comment thread is where the edge cases live
- Qwen3.8-Flash-Next model card — 158K+ downloads at 4B likes as of Aug 31; the 125B-A6B hybrid Gated DeltaNet MoE that slotstream targets
- MEASUREMENTS.md in the slotstream repo — every number ships with how it was measured, including the failed experiments
- docs/HARDWARE.md — ten-minute procedure to measure your own SSD and add your row to the table
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.