Fine-Tuning & Adaptation · LLM Reasoning · Efficient AI

On-Policy Distillation Explained: Learning From Your Own Mistakes

On-policy distillation trains a small student on its own rollouts, scored by a teacher. Five 2026 papers show how to stabilize it, prune it to 5% of tokens, and stretch it to speculative decoding and image models.

On-Policy Distillation Explained: Learning From Your Own Mistakes

What on-policy distillation actually is

Classic knowledge distillation trains a small student to imitate a large teacher on a fixed dataset. That is off-policy: the student learns from text it never generated. On-policy distillation (OPD) flips the data source. The student samples its own rollouts, and the teacher scores every token of what the student actually produced. The objective is still KL divergence, pushing the student’s token distribution toward the teacher’s, but the context is the student’s own trajectory, not a canned corpus.

The appeal is that the student learns to recover from its own mistakes, on the state distribution it will actually visit at inference. The cost is that those same mistakes can produce brutal gradients. Every paper in this guide is, in one way or another, an answer to that single problem.

OPD sits near the intersection of distillation and RL, but it is neither. On the Geometry of On-Policy Distillation measured the parameter-space footprint of OPD, SFT, and RLVR on the same setup and found OPD outside the SFT–RLVR line: its updates touch fewer weights than SFT, avoid principal directions more strongly, and lock into a narrow low-dimensional subspace early in training. Constraining later updates to that early subspace preserves OPD’s performance while degrading SFT under the same constraint. So “OPD is RL-flavored SFT” is wrong; it is its own regime with its own dynamics.

How it works

The training loop has three steps, repeated:

  1. Rollout. The student generates a continuation from a prompt, or in non-language settings its own denoising trajectory or drafting states.
  2. Scoring. The teacher computes a token-level distribution over the same positions. Standard OPD minimizes reverse KL: it sharply pushes the student toward the teacher’s most likely tokens, which is exactly what makes hard positions volatile.
  3. Update. Gradients flow on every token, including the ones where the student has drifted far from the teacher. That is where the two failure families below come from.

The first family is gradient trust: when the student generates a token the teacher considers very unlikely, reverse-KL fires a large, noisy gradient precisely where the student is weakest, and OPD destabilizes on the hard, exploratory tokens that reasoning depends on. The second family is distribution mismatch: even when gradients are well-behaved, the states the student visits at inference may simply never appear in offline training data: the draft model that never sees its own rejected proposals, the image model that never visits states produced under conflicting rewards. The four method papers below attack one family or the other.

The trust problem: TrOPD vs TA-OPD

Both papers ask the same question, namely on which tokens the teacher should be allowed to drive the update, and answer it differently.

TrOPD (Trust Region On-Policy Distillation) uses a hard gate. Each token gets a trust ratio: teacher probability divided by student probability, clipped at 1. Inside the trust region (teacher mass at least student mass), standard OPD applies. Outside it, the unstable reverse-KL gradient is replaced by a forward-KL objective over the teacher’s top-k vocabulary, plus a small off-policy warmup (beta = 0.001, annealed to zero) where the student imitates teacher-written prefixes. The result holds across setups: +3.06 average on single-domain math (49.85 vs 46.79 for OPD, AIME 2024 at 38.54 vs 35.83), +3.52 on multi-domain with a DeepSeek-Qwen-1.5B student (40.63 vs 37.11, GPQA 36.24 vs 28.03), +3.44 with a Qwen3-SFT-1.7B student. It beats all three OPD baselines tested: OPD, EOPD, and REOPOLD.

Token Teachability uses a softer filter backed by a diagnostic. The authors split per-token disagreement into learnable disagreement (teacher mass landing inside the student’s top-K candidates, a correction the student can absorb in one step) and incompatible disagreement (teacher mass outside that support), which inflates KL but yields little usable gradient. Their Teachability-Aware OPD (TA-OPD) supervises only the top tokens by a teachability score (normalized disagreement times normalized compatibility). At a 10% supervised-token budget it beats full-token OPD by +2.52 average on Qwen3-4B to Qwen3-1.7B (44.89 vs 42.37 across AIME24/25, GPQA-Diamond, HumanEval, IFEval, MATH-500), and by +3.16 on Qwen3-8B-GRPO to Qwen3-4B (56.87 vs 53.71). On the 8B-to-4B pair, the 5% budget scores 57.35, beating the 10%, 30%, and 50% budgets outright.

TrOPD and TA-OPD were never run against each other. TrOPD’s gains come from changing the objective on untrusted tokens; TA-OPD’s come from skipping tokens that carry no learnable signal. They are compatible ideas, and no paper in this set has tested the combination. TA-OPD’s own ablation shows that mixing entropy into its selector hurts (42.32 vs 44.89 clean), so the ordering of filters is not obvious.

One caveat TA-OPD states openly and the TrOPD abstract understates: these are supervision budgets, not speedups. Both teacher and student still run a forward pass over the full sequence, so a 5% token budget is not a 20x training accelerator.

The mismatch problem: Draft-OPD and Flow-OPD

The second family of fixes targets places where the gap between training data and inference states is structural, not incidental.

Draft-OPD applies OPD to speculative decoding. Draft models are normally SFT-trained on the target model’s transcripts, offline data that never shows the states the draft actually visits once it starts proposing tokens autoregressively. Accepted length plateaus because the draft is good at predicting the target’s text but bad at recovering from its own rejected proposals. Draft-OPD collects target-assisted rollouts through real speculative decoding, replays drafting from each anchor position including rejected tokens, and trains with an asymmetric KL: forward KL on accepted tokens, reverse KL on rejected ones, weighted to emphasize early rejections. On Qwen3 thinking models it reaches 4.86x speedup on Qwen3-4B and 4.89x on Qwen3-8B, versus 3.87x / 4.06x for EAGLE-3 and 4.33x / 4.34x for DFlash, roughly +23% over EAGLE-3 and +13% over DFlash under matched compute, with lossless decoding by construction.

Flow-OPD applies OPD to text-to-image models, where the analogous failure is reward conflict. Running GRPO on Stable Diffusion 3.5 with several rewards summed at once makes the objectives fight: pushing OCR drags down prompt alignment, pushing aesthetics invites reward hacking. Flow-OPD trains one GRPO specialist teacher per reward, then distills them all on-policy into a single student. The student generates its own denoising trajectories, and each teacher supplies dense per-timestep supervision on the predicted vector field, far more informative than a scalar reward that only arrives at the end of generation. A Manifold Anchor Regularization term, a time-weighted L2 penalty against a frozen quality-tuned teacher, keeps the student on the natural-image manifold. The merged student lifts GenEval from 0.63 to 0.92 and OCR accuracy from 0.59 to 0.94, while DeQA (4.35 vs 4.07) and PickScore (23.08 vs 21.64) rise rather than collapse. On the combined reward curve, the student reaches roughly 93 where vanilla multi-reward GRPO stalls near 79.

Key numbers

DomainMethodHeadline numberBaseline it beatsSource
LLM reasoningTrOPD+3.06 to +3.52 avg over OPDOPD, EOPD, REOPOLDTrOPD
LLM reasoningTA-OPD+2.52 avg at 10% token budget; 5% budget best on 8B→4BFull-token OPD, entropy selection, TIPToken Teachability
Speculative decodingDraft-OPD4.86x / 4.89x speedupEAGLE-3 3.87x / 4.06x; DFlash 4.33x / 4.34xDraft-OPD
Text-to-image flowFlow-OPDGenEval 0.63→0.92, OCR 0.59→0.94Multi-reward GRPO stalls ~79 on combined rewardFlow-OPD
DiagnosticsGeometry of OPDSubspace locking: early low-rank channel is functionally sufficientSFT degrades under the same constraintGeometry

Cross-cutting cautions the table hides. The TrOPD and TA-OPD rows come from separate harnesses (students differ: DeepSeek-Qwen2.5-1.5B vs Qwen3-1.7B; teachers differ; benchmarks differ), so their +3-ish margins are not comparable to each other, only to their own baselines. The TA-OPD 14B→4B result is a tie (54.65 vs 54.64), which says the teachability margin shrinks as the teacher-student gap narrows. And the Draft-OPD speedups are on Qwen3 thinking models at temperature 0; the headline “past 5x” aggregates across benchmarks and modes, so the per-model cells above are the cleaner comparison.

What the geometry paper implies for practitioners

The geometry study proposes no method, but it changes how you should read the other four. OPD’s cumulative updates enter a narrow low-dimensional subspace early in training and stay there; sparsifying which tokens get updated does not disturb this; shifting rollouts off-policy does not either; but mixing the OPD objective with RLVR does. Two practical corollaries follow.

First, token-selection methods like TA-OPD and TrOPD are probably safe to combine with other training changes, since the rank dynamics are robust to exactly that kind of intervention. Second, the popular recipe of interleaving OPD with RLVR to “distill then sharpen” may carry a hidden cost, because the mixture perturbs the geometry that makes OPD sample-efficient. The paper stops short of demonstrating an engineering payoff from the locked subspace, so treat low-rank OPD adapters as a hypothesis, not a recipe.

When to use which

  • Compressing a reasoning teacher into a small student, and your OPD runs destabilize on hard tokens: start with TrOPD. The gate is cheap, interpretable, and worth about 3 average points at 1.5B–1.7B scale.
  • Your OPD runs are stable but you suspect most of the KL signal is wasted: TA-OPD-style teachability filtering buys 2–3 points at a fraction of the supervised tokens. Do not expect a wall-clock speedup.
  • You operate an inference stack with speculative decoding: Draft-OPD is a drop-in upgrade over SFT-trained drafts, with +23% speedup over EAGLE-3 on Qwen3 thinking models and unchanged output quality.
  • You fine-tune flow-matching image models against multiple rewards: decompose into specialist teachers and distill on-policy, as Flow-OPD does, rather than summing rewards into one GRPO run.
  • Choosing between SFT, OPD, and RLVR from scratch: the geometry paper says the choice is not a dial between two poles. OPD is its own regime; pick it when you need on-trajectory learning, RLVR when you need exact-reward optimization, SFT when you need neither.

Limits and open questions

The honest gaps, aggregated across all five papers. Every measured gain is at small or moderate scale: TrOPD’s students are 1.5B–1.7B with 4B–7B teachers, TA-OPD’s biggest teacher is 14B, and none of them tests whether the margins survive at frontier scale or across a much wider teacher-student gap; TA-OPD’s own 14B→4B result ties at +0.01. Evaluation is heavily math-and-code flavored; multilingual, dialogue, and open-ended generation are barely touched. The two token-selection methods were never compared or combined under one harness. TrOPD’s gains are never decomposed into trust-mask versus outlier-forward-KL versus warmup contributions. And the geometry findings, while robust within the studied setup, are correlational signatures: they describe where OPD moves in weight space, not a mechanistic reason it generalizes.

FAQ

How does on-policy distillation work?

The student generates its own rollouts, the teacher scores every token with its own distribution, and the student is updated with a KL objective against the teacher on those positions, but on the student’s trajectories, not a fixed dataset. That is the difference from classic (off-policy) distillation, and from SFT, which trains on external data.

What results does TrOPD report over standard OPD?

+3.06 average on single-domain math (49.85 vs 46.79, with AIME 2024 at 38.54 vs 35.83), +3.52 on a multi-domain setup with a DeepSeek-Qwen-1.5B student (40.63 vs 37.11), and +3.44 with a Qwen3-SFT-1.7B student, beating OPD, EOPD, and REOPOLD in every reported setup.

How can TA-OPD match full-token distillation while supervising only 5% of tokens?

Because most token-level KL is incompatible disagreement: teacher mass sitting outside the student’s top-K candidates, which inflates the divergence but produces little usable gradient. Selecting the learnable slice (teacher mass inside the student’s support) concentrates supervision where one gradient step can actually absorb it; on Qwen3-8B-GRPO to Qwen3-4B the 5% budget even beats the 10–50% budgets at 57.35.

Does on-policy distillation need a reward model?

No. That is one of its selling points relative to RLVR: the teacher’s token-level distribution plays the role of dense supervision, so no verifiable reward or reward model is required. Flow-OPD is the exception that uses rewards, but it uses them to train specialist GRPO teachers, then replaces reward optimization with distillation in the second stage.

What are the limitations of on-policy distillation as of 2026?

All reported margins are at 1.5B–8B student scale with moderate teachers, so frontier-scale behavior is unproven; token-selection gains can collapse to a tie as the teacher-student gap narrows; supervision budgets do not translate into wall-clock speedups; and the update geometry that makes OPD efficient is disturbed if you mix the OPD objective with RLVR in the same run.

One line: on-policy distillation works because the student learns from its own mistakes, and the last year of work has been about deciding, token by token, which mistakes the teacher is actually qualified to grade. Read the papers: TrOPD, Token Teachability, Geometry of OPD, Draft-OPD, Flow-OPD.