Spark-X2.5: XHToken's New-Architecture SLM Beats Qwen3.5-9B on τ³-Bench, AIME 2026, and BrowseComp at 1.7B / 4B — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Spark-X2.5: XHToken's New-Architecture SLM Beats Qwen3.5-9B on τ³-Bench, AIME 2026, and BrowseComp at 1.7B / 4B

XHToken shipped Spark-X2.5-1.7B and -4B on Sep 1 with their own architecture (model_type=spark2_5) — hybrid 1 full-attention + 3 sliding-window layers, 20T tokens pretraining, 1M-token context, integrated with Codex/Claude Code/OpenClaw/Hermes. Apache-2.0, day-0 across vLLM/SGLang/llama.cpp/MLX/Ollama.

The “new SLM” stories in 2026 have mostly been fine-tunes of Qwen or Gemma — same architectures, different SFT data, marginal benchmark wins. Spark-X2.5 from XHToken is the rare case where someone shipped a new architecture family in the SLM class. Both sizes (1.7B and 4B) carry architectures = ["Spark2_5ForCausalLM"] and model_type = "spark2_5" in the HF config — not a fine-tune, not a LoRA, not a merge. Apache-2.0, day-0 across vLLM / SGLang / llama.cpp / MLX / Ollama / LM Studio. As of Sep 3 the 1.7B is at 1,128 downloads + 52 likes.

The interesting design choice isn’t the architecture itself — hybrid attention has been done. It’s the ratio: one full-attention layer for every three sliding-window attention layers. That ratio decides how far each token “sees” in the long-context regime, and most prior hybrid-attention work picked ratios like 1:5 or 1:8 to save more FLOPs. 1:3 is the small-model sweet spot where you keep enough global context for tool-use and agentic workflows without paying the full quadratic-attention tax.

The architecture in one paragraph

Each transformer block is one of two types: a full self-attention layer, or a sliding-window attention (SWA) layer with a fixed window. The block pattern is SWA, SWA, SWA, FULL, SWA, SWA, SWA, FULL, ... — one full layer per three SWA layers. SWA layers look back W tokens (the model card doesn’t say what W is, but the standard pick for the 1.7B / 4B class is 4,096 or 8,192). Full layers see the entire context up to the 1M-token limit.

The pretraining runs to 20T tokens. Long-context training is a separate stage with hundreds of billions of tokens at sequence lengths extending to 1M. Post-training is a two-stage pipeline: supervised fine-tuning on a curated corpus, then domain-specialized RL producing teacher policies that get merged via MOPD (Multi-Objective Policy Distillation) into the final deployable model. Training hardware is Huawei Ascend clusters.

The integration list reads like a roll call of where 2026 agent workflows actually run: Codex, Claude Code, OpenClaw, Hermes. That’s a deliberate choice — agent harness compatibility is treated as a first-class capability, not a side-effect of fine-tuning.

The benchmark numbers that matter

The full table in the model card compares against Qwen3.5-2B / 4B / 9B and Gemma4-E2B / E4B / 12B. Spark-X2.5-4B wins outright on τ³-bench (30.4 vs Qwen3.5-9B’s 9.3, Gemma4-12B’s 13.3), MCP-Atlas (54.6 vs 47.4 / 30.5), MCP-Mark (14.2 vs 13.4 / –), Workspace Bench (31.2 vs 25.5 / –), VitaBench 2.0 (25.2 vs 15.6 / 12.4), BrowseComp (40.9 vs 8.3 / 10.0), SWE-Bench Pro (44.4 vs 33.8 / 21.9), SWE-Bench Multilingual (53.3 vs 43.3 / 32.5), AIME 2026 (90.7 vs 88.2 / 82.1), HMMT Feb 2026 (81.2 vs 70.8 / 65.6), IMO-AnswerBench (74.2 vs 69.8 / 57.2), and IFBench (75.0 vs 64.5 / 73.5).

The 1.7B variant wins on its own tier — τ³-bench 20.1 vs Qwen3.5-2B’s 4.1, VitaBench 8.3 vs 5.2, BrowseComp 29.7 vs 3.1, AIME 2026 69.4 vs 30.8, HMMT 48.4 vs 21.5. The pattern is consistent: agent / tool-use / long-context reasoning benchmarks favor the hybrid architecture, raw knowledge benchmarks (GPQA, HLE) still favor Qwen3.5-9B.

BenchmarkSpark-X2.5-4BSpark-X2.5-1.7BQwen3.5-9BGemma4-12B
τ³-bench30.420.19.313.3
MCP-Atlas54.623.447.430.5
BrowseComp40.929.78.310.0
SWE-Bench Pro44.410.433.821.9
AIME 202690.769.488.282.1
GPQA67.443.877.272.8
HLE12.36.314.313.1

All evaluations in thinking mode. Sampling parameters: temperature=1.0, top_p=0.95, top_k=-1.

What this signals

The SLM class in 2026 is bifurcating into two camps: vendor fine-tunes of Qwen3.5 / Gemma4, and from-scratch new architectures. The fine-tunes win on raw knowledge benchmarks where the base model already has the data; the new architectures win on agent / tool-use / long-context where the architectural choice matters more than the pretraining corpus. Spark-X2.5 is the clearest example of the second camp so far.

The training-hardware signal is also worth noting: Spark-X2.5 was trained on Huawei Ascend clusters, not NVIDIA. The same was true of GLM-5.3-Flash and Tencent Hy4 — three of the past month’s biggest open-weight releases all shipped from Chinese hardware. The “frontier-class model on commodity hardware” story that slotstream pushed yesterday (Sep 1) and the “frontier-class model trained on Chinese hardware” story that Spark-X2.5 pushed today are part of the same decoupling: neither axis is GPU-vendor-locked any more.

The OpenClaw / Hermes integration is the line I want to flag specifically. Most agent-harness compatibility stories are “we tested against Claude Code once and it worked.” The Spark-X2.5 model card lists Codex / Claude Code / OpenClaw / Hermes as deeply integrated, with model-specific optimizations in the harness. That’s the difference between “drop-in compatible” and “tuned for the workflow” — and for someone running an OpenClaw-based agent pipeline, the tuned path matters.

Trade-offs and what to verify before deploying

The training corpus is a black box. 20T tokens is a serious pretraining run, but XHToken doesn’t publish the data mixture, and the licensing is Apache-2.0 for the weights but not the data recipe. If you’re training downstream or doing continual learning, you can’t reproduce the pretraining distribution.

The 1M-token context claim is a “native” claim, not a “useful” claim. Most long-context benchmarks in the model card cluster around 1K-100K tokens (BrowseComp, MCP-Atlas, τ³-bench, SWE-Bench). Whether the model is actually strong at 100K-1M tokens on agent-style tasks is a separate measurement that requires needle-in-a-haystack evaluation, multi-hop reasoning at scale, and tool-use over very long histories. The model card doesn’t publish those numbers as of Sep 3.

The 1.7B variant on SWE-Bench Pro is 10.4. That’s noticeably lower than Qwen3.5-9B’s 33.8 — 4x larger model, 3.3x higher score on hard engineering tasks. If you’re deploying for agentic coding at scale, the 4B variant is the floor, not the 1.7B.

The fine-tune ecosystem hasn’t caught up. Custom code is required to load (trust_remote_code=True), and vLLM/SGLang support landed day-0 but llama.cpp’s ggml-org/llama.cpp repo support for model_type=spark2_5 is going to lag by 1-2 weeks. ExLlamav3 is the same. If your inference path runs through llama.cpp and you can’t upgrade, wait.

Where to dig further

The training pipeline in more detail

The post-training recipe is the part most SLM release announcements skip over, and it’s the part that distinguishes Spark-X2.5 from a typical fine-tune. The first stage is supervised fine-tuning on a curated corpus — standard practice, but the model card flags that the SFT corpus is hand-picked for “robust instruction following, structured generation, and task-completion, while providing a stable policy initialization for reinforcement learning.” That’s a careful framing: the SFT stage isn’t trying to make the model good at chat, it’s trying to make the model trainable in the RL stage that follows.

The RL stage is multi-domain. Separate teacher policies get trained on language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following. Each teacher policy is domain-specialized — they aren’t generalists trained to be slightly better at everything, they’re specialists that excel in one area. The final deployable model gets assembled by MOPD (Multi-Objective Policy Distillation), which is the part that’s actually novel.

MOPD is the consolidation step. Multiple teacher policies have complementary strengths but can’t be deployed as a mixture-of-experts or a router-fronted ensemble — that would require runtime routing and would blow the inference budget. So the teacher policies get distilled into a single model, with the multi-objective optimization preserving each specialist’s strength in its area. The model card describes this as “domain-specialized teacher policies whose complementary strengths are consolidated into a single deployable model.” That sentence is doing a lot of work — it’s the difference between “five fine-tunes merged with averaging” and a learned consolidation that picks the right behavior per task.

Practically: this is why Spark-X2.5-4B wins on τ³-bench (multi-turn tool use) and AIME 2026 (math reasoning) and SWE-Bench Pro (engineering) without collapsing into a “good at chat, mediocre at everything else” pattern. The distillation preserves the specialists. Whether MOPD is genuinely new or a rebrand of an existing multi-task distillation technique isn’t clear from the public docs — the XHToken team hasn’t published a paper yet. If MOPD is novel, expect a paper in Q4 2026. If it’s a rebrand, the model card’s framing of “domain-specialized teacher policies consolidated into a single deployable model” is the recipe, and other teams can copy it.

The deployment realities

The day-0 framework support list is what makes Spark-X2.5 ship-able now rather than as a research curiosity. vLLM and SGLang picked it up at release. llama.cpp support for model_type=spark2_5 is the question — the GGML community typically needs 1-2 weeks after a new architecture to land native kernels, and custom-code paths require explicit support. MLX support landed for Apple Silicon, and Ollama / LM Studio can wrap it.

The custom-code requirement is real. The HF config has auto_map entries for AutoConfig, AutoModel, and AutoModelForCausalLM, all pointing at custom code in the model repo. Loading without trust_remote_code=True will fail. That’s fine for a single-machine deploy where you control the trust boundary; it’s a non-starter for production inference where you can’t pass arbitrary Python code to the loader. Expect a static-graph rewrite or a converter that drops the custom-code path within 2-3 weeks of release, once the vLLM/SGLang teams have had time to study the architecture.

The hardware compatibility list is unusual: NVIDIA, Huawei, Hygon, HOUMO.AI. Hygon is the Chinese x86 server CPU line that came out of the AMD licensing partnership; HOUMO.AI is a domestic Chinese GPU vendor. The model was trained on Huawei Ascend, so the inference support on Ascend hardware is presumably first-class. If you’re running a Chinese-vendor inference stack (which a non-trivial fraction of 2026 enterprise deployments are), Spark-X2.5 is one of the few new SLM releases with native support across the vendor matrix.

What I’d want to see in the next release

Three things would make Spark-X2.5 the default SLM pick in my workflow: (1) a static-graph rewrite that drops the trust_remote_code requirement so I can deploy it in production without trusting arbitrary Python, (2) llama.cpp support for model_type=spark2_5 so I can run it on Apple Silicon without the MLX wrapper, (3) needle-in-a-haystack and multi-hop reasoning benchmarks at the 100K-1M context length so I know whether the “native 1M context” claim is useful or just marketing. The base checkpoints (XHToken/Spark-X2.5-1.7B-Base, XHToken/Spark-X2.5-4B-Base) are available for further fine-tuning, which makes the 4B variant a serious candidate for domain-specialized downstream training if you have the compute.

The 30.4 on τ³-bench at 4B parameters is the single number I’d watch. That’s a tool-use benchmark designed to test multi-turn agent behavior, and Qwen3.5-9B sits at 9.3 on the same metric. Either the benchmark is suddenly saturating (unlikely — τ³-bench was released in late 2025 specifically because τ² was saturating), or the hybrid 1:3 attention ratio is genuinely better for tool-use than pure attention at this scale. I’m betting on the second — but it’s worth waiting for the reproducibility reports before locking it in.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.