The “new SLM” stories in 2026 have mostly been fine-tunes of Qwen or Gemma — same architectures, different SFT data, marginal benchmark wins. Spark-X2.5 from XHToken is the rare case where someone shipped a new architecture family in the SLM class. Both sizes (1.7B and 4B) carry architectures = ["Spark2_5ForCausalLM"] and model_type = "spark2_5" in the HF config — not a fine-tune, not a LoRA, not a merge. Apache-2.0, day-0 across vLLM / SGLang / llama.cpp / MLX / Ollama / LM Studio. As of Sep 3 the 1.7B is at 1,128 downloads + 52 likes.
The interesting design choice isn’t the architecture itself — hybrid attention has been done. It’s the ratio: one full-attention layer for every three sliding-window attention layers. That ratio decides how far each token “sees” in the long-context regime, and most prior hybrid-attention work picked ratios like 1:5 or 1:8 to save more FLOPs. 1:3 is the small-model sweet spot where you keep enough global context for tool-use and agentic workflows without paying the full quadratic-attention tax.
The architecture in one paragraph
Each transformer block is one of two types: a full self-attention layer, or a sliding-window attention (SWA) layer with a fixed window. The block pattern is SWA, SWA, SWA, FULL, SWA, SWA, SWA, FULL, ... — one full layer per three SWA layers. SWA layers look back W tokens (the model card doesn’t say what W is, but the standard pick for the 1.7B / 4B class is 4,096 or 8,192). Full layers see the entire context up to the 1M-token limit.
The pretraining runs to 20T tokens. Long-context training is a separate stage with hundreds of billions of tokens at sequence lengths extending to 1M. Post-training is a two-stage pipeline: supervised fine-tuning on a curated corpus, then domain-specialized RL producing teacher policies that get merged via MOPD (Multi-Objective Policy Distillation) into the final deployable model. Training hardware is Huawei Ascend clusters.
The integration list reads like a roll call of where 2026 agent workflows actually run: Codex, Claude Code, OpenClaw, Hermes. That’s a deliberate choice — agent harness compatibility is treated as a first-class capability, not a side-effect of fine-tuning.
The benchmark numbers that matter
The full table in the model card compares against Qwen3.5-2B / 4B / 9B and Gemma4-E2B / E4B / 12B. Spark-X2.5-4B wins outright on τ³-bench (30.4 vs Qwen3.5-9B’s 9.3, Gemma4-12B’s 13.3), MCP-Atlas (54.6 vs 47.4 / 30.5), MCP-Mark (14.2 vs 13.4 / –), Workspace Bench (31.2 vs 25.5 / –), VitaBench 2.0 (25.2 vs 15.6 / 12.4), BrowseComp (40.9 vs 8.3 / 10.0), SWE-Bench Pro (44.4 vs 33.8 / 21.9), SWE-Bench Multilingual (53.3 vs 43.3 / 32.5), AIME 2026 (90.7 vs 88.2 / 82.1), HMMT Feb 2026 (81.2 vs 70.8 / 65.6), IMO-AnswerBench (74.2 vs 69.8 / 57.2), and IFBench (75.0 vs 64.5 / 73.5).
The 1.7B variant wins on its own tier — τ³-bench 20.1 vs Qwen3.5-2B’s 4.1, VitaBench 8.3 vs 5.2, BrowseComp 29.7 vs 3.1, AIME 2026 69.4 vs 30.8, HMMT 48.4 vs 21.5. The pattern is consistent: agent / tool-use / long-context reasoning benchmarks favor the hybrid architecture, raw knowledge benchmarks (GPQA, HLE) still favor Qwen3.5-9B.
| Benchmark | Spark-X2.5-4B | Spark-X2.5-1.7B | Qwen3.5-9B | Gemma4-12B |
|---|---|---|---|---|
| τ³-bench | 30.4 | 20.1 | 9.3 | 13.3 |
| MCP-Atlas | 54.6 | 23.4 | 47.4 | 30.5 |
| BrowseComp | 40.9 | 29.7 | 8.3 | 10.0 |
| SWE-Bench Pro | 44.4 | 10.4 | 33.8 | 21.9 |
| AIME 2026 | 90.7 | 69.4 | 88.2 | 82.1 |
| GPQA | 67.4 | 43.8 | 77.2 | 72.8 |
| HLE | 12.3 | 6.3 | 14.3 | 13.1 |
All evaluations in thinking mode. Sampling parameters: temperature=1.0, top_p=0.95, top_k=-1.
What this signals
The SLM class in 2026 is bifurcating into two camps: vendor fine-tunes of Qwen3.5 / Gemma4, and from-scratch new architectures. The fine-tunes win on raw knowledge benchmarks where the base model already has the data; the new architectures win on agent / tool-use / long-context where the architectural choice matters more than the pretraining corpus. Spark-X2.5 is the clearest example of the second camp so far.
The training-hardware signal is also worth noting: Spark-X2.5 was trained on Huawei Ascend clusters, not NVIDIA. The same was true of GLM-5.3-Flash and Tencent Hy4 — three of the past month’s biggest open-weight releases all shipped from Chinese hardware. The “frontier-class model on commodity hardware” story that slotstream pushed yesterday (Sep 1) and the “frontier-class model trained on Chinese hardware” story that Spark-X2.5 pushed today are part of the same decoupling: neither axis is GPU-vendor-locked any more.
The OpenClaw / Hermes integration is the line I want to flag specifically. Most agent-harness compatibility stories are “we tested against Claude Code once and it worked.” The Spark-X2.5 model card lists Codex / Claude Code / OpenClaw / Hermes as deeply integrated, with model-specific optimizations in the harness. That’s the difference between “drop-in compatible” and “tuned for the workflow” — and for someone running an OpenClaw-based agent pipeline, the tuned path matters.
Trade-offs and what to verify before deploying
The training corpus is a black box. 20T tokens is a serious pretraining run, but XHToken doesn’t publish the data mixture, and the licensing is Apache-2.0 for the weights but not the data recipe. If you’re training downstream or doing continual learning, you can’t reproduce the pretraining distribution.
The 1M-token context claim is a “native” claim, not a “useful” claim. Most long-context benchmarks in the model card cluster around 1K-100K tokens (BrowseComp, MCP-Atlas, τ³-bench, SWE-Bench). Whether the model is actually strong at 100K-1M tokens on agent-style tasks is a separate measurement that requires needle-in-a-haystack evaluation, multi-hop reasoning at scale, and tool-use over very long histories. The model card doesn’t publish those numbers as of Sep 3.
The 1.7B variant on SWE-Bench Pro is 10.4. That’s noticeably lower than Qwen3.5-9B’s 33.8 — 4x larger model, 3.3x higher score on hard engineering tasks. If you’re deploying for agentic coding at scale, the 4B variant is the floor, not the 1.7B.
The fine-tune ecosystem hasn’t caught up. Custom code is required to load (trust_remote_code=True), and vLLM/SGLang support landed day-0 but llama.cpp’s ggml-org/llama.cpp repo support for model_type=spark2_5 is going to lag by 1-2 weeks. ExLlamav3 is the same. If your inference path runs through llama.cpp and you can’t upgrade, wait.
Where to dig further
- Spark-X2.5-1.7B on Hugging Face — model card, Apache-2.0, config + custom code
- Spark-X2.5-4B on Hugging Face — larger variant, same architecture
- r/LocalLLaMA release thread — community benchmarks and quirks
- XHToken organization — base checkpoints, longer-context variants, MOPD tooling
- SWE-Bench Pro paper — the saturation-resistant SWE benchmark where Spark-X2.5-4B’s 44.4 lands
The training pipeline in more detail
The post-training recipe is the part most SLM release announcements skip over, and it’s the part that distinguishes Spark-X2.5 from a typical fine-tune. The first stage is supervised fine-tuning on a curated corpus — standard practice, but the model card flags that the SFT corpus is hand-picked for “robust instruction following, structured generation, and task-completion, while providing a stable policy initialization for reinforcement learning.” That’s a careful framing: the SFT stage isn’t trying to make the model good at chat, it’s trying to make the model trainable in the RL stage that follows.
The RL stage is multi-domain. Separate teacher policies get trained on language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following. Each teacher policy is domain-specialized — they aren’t generalists trained to be slightly better at everything, they’re specialists that excel in one area. The final deployable model gets assembled by MOPD (Multi-Objective Policy Distillation), which is the part that’s actually novel.
MOPD is the consolidation step. Multiple teacher policies have complementary strengths but can’t be deployed as a mixture-of-experts or a router-fronted ensemble — that would require runtime routing and would blow the inference budget. So the teacher policies get distilled into a single model, with the multi-objective optimization preserving each specialist’s strength in its area. The model card describes this as “domain-specialized teacher policies whose complementary strengths are consolidated into a single deployable model.” That sentence is doing a lot of work — it’s the difference between “five fine-tunes merged with averaging” and a learned consolidation that picks the right behavior per task.
Practically: this is why Spark-X2.5-4B wins on τ³-bench (multi-turn tool use) and AIME 2026 (math reasoning) and SWE-Bench Pro (engineering) without collapsing into a “good at chat, mediocre at everything else” pattern. The distillation preserves the specialists. Whether MOPD is genuinely new or a rebrand of an existing multi-task distillation technique isn’t clear from the public docs — the XHToken team hasn’t published a paper yet. If MOPD is novel, expect a paper in Q4 2026. If it’s a rebrand, the model card’s framing of “domain-specialized teacher policies consolidated into a single deployable model” is the recipe, and other teams can copy it.
The deployment realities
The day-0 framework support list is what makes Spark-X2.5 ship-able now rather than as a research curiosity. vLLM and SGLang picked it up at release. llama.cpp support for model_type=spark2_5 is the question — the GGML community typically needs 1-2 weeks after a new architecture to land native kernels, and custom-code paths require explicit support. MLX support landed for Apple Silicon, and Ollama / LM Studio can wrap it.
The custom-code requirement is real. The HF config has auto_map entries for AutoConfig, AutoModel, and AutoModelForCausalLM, all pointing at custom code in the model repo. Loading without trust_remote_code=True will fail. That’s fine for a single-machine deploy where you control the trust boundary; it’s a non-starter for production inference where you can’t pass arbitrary Python code to the loader. Expect a static-graph rewrite or a converter that drops the custom-code path within 2-3 weeks of release, once the vLLM/SGLang teams have had time to study the architecture.
The hardware compatibility list is unusual: NVIDIA, Huawei, Hygon, HOUMO.AI. Hygon is the Chinese x86 server CPU line that came out of the AMD licensing partnership; HOUMO.AI is a domestic Chinese GPU vendor. The model was trained on Huawei Ascend, so the inference support on Ascend hardware is presumably first-class. If you’re running a Chinese-vendor inference stack (which a non-trivial fraction of 2026 enterprise deployments are), Spark-X2.5 is one of the few new SLM releases with native support across the vendor matrix.
What I’d want to see in the next release
Three things would make Spark-X2.5 the default SLM pick in my workflow: (1) a static-graph rewrite that drops the trust_remote_code requirement so I can deploy it in production without trusting arbitrary Python, (2) llama.cpp support for model_type=spark2_5 so I can run it on Apple Silicon without the MLX wrapper, (3) needle-in-a-haystack and multi-hop reasoning benchmarks at the 100K-1M context length so I know whether the “native 1M context” claim is useful or just marketing. The base checkpoints (XHToken/Spark-X2.5-1.7B-Base, XHToken/Spark-X2.5-4B-Base) are available for further fine-tuning, which makes the 4B variant a serious candidate for domain-specialized downstream training if you have the compute.
The 30.4 on τ³-bench at 4B parameters is the single number I’d watch. That’s a tool-use benchmark designed to test multi-turn agent behavior, and Qwen3.5-9B sits at 9.3 on the same metric. Either the benchmark is suddenly saturating (unlikely — τ³-bench was released in late 2025 specifically because τ² was saturating), or the hybrid 1:3 attention ratio is genuinely better for tool-use than pure attention at this scale. I’m betting on the second — but it’s worth waiting for the reproducibility reports before locking it in.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.