Miles v0.1: SGLang Rollout + Megatron Trainer, and the Agentic RL Loop That Fits 744B Into 64 GPUs — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Miles v0.1: SGLang Rollout + Megatron Trainer, and the Agentic RL Loop That Fits 744B Into 64 GPUs

RadixArk + LMSYS partners shipped Miles v0.1 on August 18, 2026 — a full-stack RL post-training system that pins SGLang as the rollout engine, lets you swap between Megatron-LM and PyTorch FSDP for the trainer, and ships verified end-to-end on a 744B-A40B GLM-5.2 agentic coding run on 64 GB300 GPUs. The interesting bit is not the framework itself but three coupled design decisions: bit-exact precision contracts between rollout and training, a session server that preserves the exact token IDs the model generated, and a fully asynchronous scheduler that decouples rollout throughput from training cadence.

The first thing you notice reading the Miles v0.1 technical report is that it is not framed as a paper about a new algorithm. It is a paper about couplings. RadixArk — the team that runs SGLang and the wider LMSYS-org RL work — published the 34-page report on arXiv on September 8, 2026 (arXiv:2609.08368), three weeks after the v0.1 release on August 18. The release itself is the second in the Miles series, successor to the original Miles framework. What changed between Miles and Miles v0.1 is, in their own words, a single principle: “components should be verified, clean, and customizable.” That sentence is doing more work than it looks like. The couplings between components — rollout to trainer, trainer to rollout, MoE routing between the two, token IDs between the engine and the loss — are exactly where frontier-scale RL breaks. The report walks through each coupling in turn, and what is interesting is that they did not invent new RL math to fix them. They engineered the interface so the math that already exists stops lying.

The framework is Apache-2.0, hosted at radixark/miles on GitHub, written in Python, with Megatron-LM and PyTorch FSDP as alternative trainer backends. As of this morning (2026-09-11) the repo had 2,797 stars, 986 open issues, and a last push six minutes before I checked. The repo description — “forked from and co-evolving with slime” — names the design lineage. Slime is the earlier RadixArk RL framework; Miles v0.1 is what you get when you take slime’s clean scheduler design and add the production plumbing that frontier post-training demands: real sandbox isolation for agentic environments, multi-backend weight transport, four precision contracts, and an end-to-end reference run on a model most teams cannot afford to train.

The RL loop, in three stages and three couplings

A Miles job is a loop over three stages: rollout (SGLang generates trajectories), training (Megatron-LM or FSDP consumes completed trajectory groups and computes the RL loss), and weight update (the trainer’s new weights are synchronized back to the rollout fleet). Each stage has a coupling to the next one, and each coupling has been a perennial source of correctness bugs in RL post-training. The report walks through all three.

The rollout-to-trainer coupling is the token stream. Multi-turn agentic RL has a property that single-turn RL does not: between turns, the model’s output passes through message parsing, tool execution, and chat-template rendering before the next turn’s input is formed. That step can re-tokenize history, prune reasoning, or reserialize tool calls — silently producing a different token context from the one the rollout engine actually consumed. The trainer would then optimize against tokens the model never generated. Miles handles this with a session server they call TITO (Token-In-Token-Out). On each new turn, TITO tokenizes only the newly appended messages and merges them with the existing prefix, preserving the exact token IDs SGLang produced. The full trajectory assembles into one contiguous training sample; the original rollout log probabilities survive; tokens that were not generated by the model get loss-masked. TITO is validated through CPU round-trip tests and real SGLang GPU sessions for each model family. This is not “we tested it once on one model.” It is a per-model contract.

The trainer-to-rollout coupling is the weight stream. Every training step produces a new policy, and the rollout fleet needs to see it. Naively, you stop the rollout, push the weights, resume. In practice on a long agentic workload the rollout fleet can spend more time paused than generating. Miles ships three weight-synchronization transports for different deployment topologies — the references list points to “Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL” — and reports that the GB300 case study below runs with rollout weights lagging on average 1.7 steps behind the trainer, with no perceptible loss of training stability. The 1.7-step lag is the cost of the asynchrony; the alternative cost, fully synchronous, is much worse.

The rollout-to-itself coupling is the MoE routing. This one is specific to mixture-of-experts models and it is the kind of bug that does not show up on dense models. A single top-k routing flip changes both the token’s computation and which expert receives the gradient. If SGLang during rollout picks a slightly different top-k than Megatron during training, the two sides disagree on which expert the gradient belongs to. Miles calls the fix R3 (Rollout Routing Replay). During rollout, R3 records SGLang’s expert routing decisions. During training, R3 replays those decisions instead of recomputing them. The replay is implemented inside SGLang so the overhead is minimal. R3 is the kind of fix that exists because someone actually saw the divergence — a MoE RL run where the reward curve looked fine for a while, then collapsed, and the only diagnosis that worked was “the experts are receiving gradients for tokens they did not route.”

The precision contracts

Low-precision RL on Blackwell needs more than swapping GEMM dtypes. Miles supports rollouts in NVFP4, MXFP4, MXFP8, and FP8, and end-to-end training recipes for NVFP4, MXFP8, and FP8. The trap is that quantization on the rollout side and the training side must agree, or the mismatch accumulates across weight updates into policy divergence. The report calls this “a bit-exact quantizer contract” — rollout kernels and training kernels must see the same quantized values, with fine-grained flags to keep sensitive layers in BF16. The MXFP8 recipe runs rollout, forward, and both gradient GEMMs with hardware blockwise scaling. The NVFP4 recipe quantizes MoE expert weights per-token with online activation scaling to avoid batch-dependent quantization artifacts. The verification is reward curves that track BF16 within tolerance while reducing rollout time. Earlier Miles posts cover the FP8 recipe in depth (“Unified FP8: Moving Beyond Mixed Precision for Stable and Accelerated MoE RL”) and the INT4 quantization-aware-training recipe (“Squeezing 1TB Model Rollout into a Single H200: INT4 QAT RL End-to-End Practice”). The August v0.1 release unified all four precisions under the same contract model.

The reference run: GLM-5.2 744B-A40B, 64 GB300, 4.5 minutes per step

The end-to-end case study is the most quotable thing in the report because it pins the abstract design choices to a concrete number on real hardware. The setup: GLM-5.2 (744B parameters, 40B active per token, MoE), trained on terminal-use coding tasks, fully asynchronous RL across 64 NVIDIA GB300 GPUs split 32 rollout + 32 training. Training parallelism is TP 2 / PP 4 / CP 4 / EP 8; inference parallelism is TP 8 / EP 8; multi-token prediction is enabled. Each multi-turn terminal agent runs in its own sandbox through the OpenEnv integration. Sequence length 65k, batch size 64. The reference run completes 100 terminal-bench-like coding task rollout steps stably, with training steps at roughly 4.5 minutes each, rollout weights lagging 1.7 steps behind the trainer on average, and a 96% prefix-cache hit rate at stable concurrency. The memory optimizations are what allow 744B to fit on 32 GB300 GPUs for training while 32 more handle rollout — optimizer states can be offloaded to CPU or node-local NVMe and streamed back per bucket during each optimizer step.

The 4.5-minute step number is the one worth holding on to. 744B-parameter agentic RL, with sandbox isolation, token-faithful training, MoE routing replay, NVFP4 rollout, and async scheduling, in 270 seconds per step, on a single 64-GPU GB300 node-aggregation. The full launch script is published; the team treats reproducibility as part of the deliverable, not a follow-up. That is unusual. Most frontier post-training systems publish numbers from a configuration you cannot recreate.

What TITO does for black-box harnesses

The piece of the report that surprised me most is the half-page on training Claude Code and Codex — agent harnesses where the model is opaque to Miles. Black-box harnesses spawn subagents and compact context at runtime; the number of trajectories per task is dynamic and unknown in advance. Miles handles this by recording the full trajectory tree through the TITO session server and applying loss normalization for varying batch sizes on the training side. The implication is that you can post-train against a closed agent system without owning the agent’s internals — the session server treats the harness as a token stream with checkpoints, and the trainer treats the result as a sequence of training samples. That is a meaningful shift from the assumption most RL post-training code makes, which is that you control the rollout loop end to end. The fact that the team chose to surface this as a TITO extension rather than a separate framework says something about how they expect the field to develop: agent harnesses will become more opaque, and post-training has to accommodate that.

Day-0 model support as a release pattern

The references list reads like a who’s-who of 2026 frontier models with a small “Day-0” tag next to each: DeepSeek-V4 (“Day-0 support — from fast inference to verified RL with SGLang and Miles”), NVIDIA Nemotron 3 Ultra (“day-0 support for long-running autonomous agents”), Inkling (“a frontier multimodal model”), Kimi K3, Qwen3.8. The pattern is that when a frontier model ships, SGLang and Miles ship first-party verified support within hours — and “verified” here means the same TITO round-trip tests and reward-curve baselines that landed with v0.1. This is a release-tempo claim worth taking seriously. If you are building agent products on top of a frontier model that releases every few weeks, the bottleneck is usually not the model weights; it is the gap between model release and the post-training stack catching up. Miles has turned that gap into a sprint.

The diversity of supported models — DeepSeek, Nemotron, Inkling, Kimi, Qwen — is also a signal. This is not a vendor’s vertically integrated stack. It is a community framework that has done the integration work for the open-weights frontier. That positions Miles differently from the closed post-training systems inside the frontier labs themselves. Whether the open-weights ecosystem will converge on a small number of such frameworks, or whether every lab keeps its own, is the open question. Miles v0.1 is a credible bet that convergence is possible.

What v0.1 does not fix

Three honest limits worth naming. First, TITO is validated for each model family individually. New model releases still require per-model CPU round-trip tests and SGLang GPU sessions before R3, OPD, and zero-KL alignment can be trusted. The cost is in the validation pass, not the framework, and it is not zero. Second, the bit-exact quantizer contract depends on SGLang and Megatron agreeing on the same quantization implementation. If a future SGLang rollout or a future Megatron release changes its low-precision kernel behavior, the contract breaks and the reward curve diverges. The mitigation is the verification harness, but the trust is per-version. Third, async RL’s 1.7-step weight lag is the floor of the asynchrony, not zero. For some loss formulations (off-policy corrections, importance sampling) that lag matters; Miles handles it with importance-ratio clipping and the bounded data buffer, but the trade-off is real and visible on the loss curve.

A note on the framework’s release history worth flagging: the repo was created on 2025-10-09 and has 986 open issues as of this morning. That issue count is high but not surprising — the framework is in active use and the community is reporting real bugs from real runs. The 2,797-star count is correspondingly modest for an Apache-2.0 framework that runs at this scale, which is itself a signal: the audience is RL post-training engineers, who do not star things lightly. The repo description’s “forked from and co-evolving with slime” sentence names the project’s relationship to its predecessor honestly. You can read the v0.1 source against slime and see what changed.

Where this leaves agentic RL

The broader question Miles v0.1 answers is whether frontier-scale agentic RL post-training can be done on open infrastructure. The August 18 release, the September 8 paper, the day-0 support list, the GLM-5.2 reference run — together they make the case that yes, with the right couplings engineered and verified, you can train a 744B agentic coding model on a 64-GPU GB300 cluster with the same rollout–trainer stack a frontier lab would use. The numbers in the report are not best-case; they are reproducible from the launch script.

Whether Miles becomes the post-training stack for the open-weights frontier or whether a different framework wins that position is a separate question. The release tempo and the day-0 support pattern are the leading indicators. Right now, the report and the GLM-5.2 case study are the clearest signal of what “production-ready” looks like for a 2026 frontier RL post-training system that does not live inside a single vendor. Aniket runs agentic pipelines that touch this surface regularly; this is the kind of system that is worth knowing the design of, even if you do not train your own 744B models on a 64-GPU cluster. The principles — bit-exact precision contracts, token-faithful training, async scheduler, MoE routing replay — are the operating constraints of any frontier RL run, regardless of whose stack you use.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.