OPSA: Distilling a 1.7B Model With No Teacher Beats On-Policy Distillation by +16.77 Points on AIME24 — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

OPSA: Distilling a 1.7B Model With No Teacher Beats On-Policy Distillation by +16.77 Points on AIME24

Purdue team's 'On-Policy Self-Adaptation' flips the SLM distillation playbook: the gains in OPD come from suppressing low-log-probability tokens, not from teacher mimicry — so a single fixed negative advantage with entropy-adaptive weights reproduces the gains without any frontier teacher. +263% relative on AIME24 vs base Qwen3-1.7B.

The SLM training playbook for the past two years has had a quiet assumption baked in: you need a frontier teacher to distill a small model into something useful. The reasoning was straightforward — small models trained on their own outputs plateau quickly, so the supervision signal has to come from somewhere larger. On-policy distillation (OPD) was the latest refinement: let the student generate trajectories, score them with the teacher, train on the dense per-token signal. Works well, costs a lot of compute, requires a frontier teacher you might not have access to.

The Purdue team’s paper Does On-Policy Distillation Really Distill? (arXiv 2608.31046, HF Daily Papers top of Sep 1 at 115↑) proposes an uncomfortable answer: OPD works largely by suppressing low log-probability tokens, which requires no teacher. The single-sentence version of the paper is “you can throw away the teacher, use a fixed negative advantage with entropy-adaptive weights, and get +263% relative on AIME24 over base Qwen3-1.7B — outperforming OPD by +16.77 points.”

If that holds up under reproducibility testing, the 2026 SLM training stack just got a lot cheaper.

What the paper actually shows

The authors started by asking whether OPD’s gains come from teacher supervision or from something else. They measured teacher noise during OPD training and found it “substantial” — the prevalence increases with teacher scale. Big teachers are noisier per-token supervisors than small teachers, which is the opposite of what you’d naively expect. The student, though, is insensitive to that noise: stripping noisy supervision matches the converged performance of keeping it.

That’s the first surprise. The second is the dissection of what OPD actually teaches the student. By analyzing which tokens get the largest gradient updates, the authors find that learning concentrates on low log-probability tokens. The teacher’s contribution, in other words, is mostly “here are tokens the student shouldn’t put probability mass on” — not “here’s what a smart model would say.” A single fixed negative advantage (no teacher, no per-token scoring) reproduces the bulk of the gain.

That motivates OPSA: On-Policy Self-Adaptation. The recipe is two ingredients. First, assign stronger learning signals to high-entropy positions — i.e. the tokens where the student is uncertain, which are exactly the low-log-probability tokens. Second, suppress those tail tokens and evenly redistribute probability mass among head tokens. No teacher, no per-token supervision signal, no frontier-model API calls.

# conceptual, not the actual algorithm — read this as the loss shape
def opsa_step(student, batch):
    logits = student(batch.tokens)
    log_probs = F.log_softmax(logits, dim=-1)
    entropy = -(log_probs * log_probs.exp()).sum(dim=-1)  # per-position entropy

    # entropy-adaptive weight: stronger signal where the student is uncertain
    weights = (entropy / entropy.max(dim=-1, keepdim=True).values)

    # fixed negative advantage: no teacher, no scoring
    advantage = -1.0 * weights  # suppress tail tokens, redistribute head mass

    # standard policy gradient on the per-token advantage
    loss = -(log_probs.gather(-1, batch.targets.unsqueeze(-1)) * advantage).mean()
    return loss

The Qwen3-1.7B base, with OPSA, gets +35.41 Avg@32 on AIME24 (263% relative gain) and more than doubles Pass@32 across all three benchmarks they tested. OPD with a real teacher gets +18.64 in the same setup. OPSA wins by +16.77.

Why this matters beyond the benchmark

The implication is not “distillation is dead” — distillation works fine, and OPSA’s claim is narrower than that. The claim is that the teacher signal in OPD is doing less than you thought, and the cheap part of OPD (suppress tail tokens) is doing more than you thought. So if you have access to a frontier teacher, OPD still works; if you don’t, OPSA gets you most of the way there at a fraction of the cost.

The cost angle is the part to focus on. OPD training requires running the teacher on every student-generated trajectory, which for a small student means paying for a large model’s inference budget at every training step. For a research lab with frontier-model access, that’s a rounding error. For a small team trying to specialize a 1.7B model for a domain, that’s a hard wall. OPSA removes the wall: the training loop is just student-on-student, no frontier model in the loop at all.

The second-order implication is about what we think we know about RLHF / RLVR. RL with verifiable rewards has been the dominant alignment-and-reasoning recipe for the past 18 months, and OPD was framed as a complementary technique for cases where the reward signal is sparse. The Purdue result suggests that the per-token supervision in OPD isn’t carrying its weight — and by extension, the per-token supervision in RLVR might be doing more mechanical work (suppressing tail tokens) and less semantic work (teaching the model what good outputs look like) than the field has assumed.

That’s a stronger claim than the paper makes explicitly. The paper is careful to say “OPD works largely by suppressing low log-probability tokens” — it doesn’t extrapolate to RLVR. But the methodological pattern is identical: per-token supervision signal, scaled by some function of position importance, applied to the student’s own outputs. If the mechanism is the same, the conclusion probably is too. Expect follow-up work.

The reproducibility question

The HF Daily Papers community score is 115↑ as of Sep 1, the repo (DripNowhy/On-Policy-Self-Adaptation) has 35 stars and a clean implementation, and the benchmark numbers are on AIME24 — a verifiable, math-reasoning benchmark with concrete ground truth. The setup is reproducible: Qwen3-1.7B base, the published hyperparameters, the standard AIME24 evaluation harness. If you’re an inference team with spare GPU time, this is a weekend project to verify the claim.

The honest thing to say: a single paper from a single team on a single benchmark family isn’t a settled result. The +35.41 on AIME24 is large enough that it’s not noise, but the absence of teacher signal is a sufficiently surprising claim that I want to see at least one independent reproduction before locking it in. The implementation being open-source is a good sign — it lowers the cost of independent verification to the point where someone will run it within a week.

The Qwen3-1.7B baseline matters here too. Qwen3-1.7B is a reasonable base, not a weak one. If OPSA’s gains came from “rescuing a model that was undertrained,” the result would be less interesting than “improving a model that’s already reasonable.” The 263% relative improvement is on top of a viable base model, which makes the method’s generalizability claim stronger.

What I’d build with this

If OPSA reproduces, the immediate application is domain-specialized SLM training without frontier-teacher access. A small team with a 7B-or-smaller base could iterate on OPSA training runs to specialize for their domain — finance, legal, code, scientific QA — without paying frontier inference costs. The cost reduction is order-of-magnitude, not incremental.

The longer-term application is what this means for the distillation industry. If the teacher signal in OPD is mostly mechanical, then the value-add of frontier-teacher APIs shifts from “per-token supervision” to “high-quality demonstrations for SFT” and “preference data for RLHF/DPO.” Those are still valuable, but they’re a smaller wedge than the distillation-everything framing suggests.

The thing I’d watch for in the next 30 days: independent reproductions on non-Qwen base models (LLaMA-3.2-1B, Gemma3-1B, Phi-3.5-mini). If OPSA generalizes across base model families, the method is real. If it’s a Qwen-specific quirk (Qwen models are known to have unusual log-prob distributions at the tail), the result is interesting but narrower than it looks.

The broader SLM training context

This paper lands at a moment when the SLM training stack was already diversifying. The dominant 2025 recipe was “fine-tune a small Qwen or LLaMA on instruction data + preference data, ship it.” The dominant 2026 recipe is something messier: hybrid attention architectures (Spark-X2.5), MoE distillation for commodity hardware (Slotstream’s target model), domain-specialized SFT (LFM2-2.6B-Longevity for biology, the various legal/finance SLMs), and now supervision-free self-improvement (OPSA). The common thread is “find a cheaper training loop that produces a model that’s competitive on the relevant benchmarks.”

OPSA fits that pattern exactly. The thing it removes from the OPD recipe — frontier-teacher inference — is the most expensive part, and the thing it adds — entropy-adaptive negative advantages — is essentially free. The Qwen3-1.7B base is small enough that the training loop runs on a single high-end consumer GPU; the published hyperparameters suggest a weekend of compute, not a week. If you’re a team that was going to train a 1-2B model anyway, the marginal cost of adding OPSA to your training pipeline is near zero.

The longer-arc question is what this implies for the frontier-lab training playbook. If a 1.7B model can get +263% on AIME24 with no teacher, the implication for 70B-and-up models is uncomfortable: their per-token supervision signals might be doing more mechanical work than the field has assumed. I don’t expect the frontier labs to publish papers questioning their own methods, but the next 6 months of independent SLM training research will probably look very different if OPSA holds up.

References and where to dig further

The honest summary: this is one paper, one team, one base model family. The mechanism it proposes — “the gains in OPD are mechanical, not semantic” — is a strong claim. If it holds up, the SLM training playbook for teams without frontier-teacher access changes substantially. If it doesn’t, it’s still a useful correction to the assumption that OPD gains come from teacher mimicry. Either way, it’s a paper worth reading carefully.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.