how much RL-acquired capability transfers across model scales via on-policy distillation, how fast, and how to predict it before training.
On-policy distillation (OPD) from a small RL-trained expert reliably lifts much larger students, and the outcome is predictable before training: gold score first rises linearly in $\sqrt{\mathrm{KL}}$, and the peak follows a power law in student size, teacher size, and teacher score.
Reinforcement learning (RL) instills strong reasoning skills in LLMs, but it is expensive to repeat at every model size. On-policy distillation offers a shortcut: a teacher scores every token of rollouts sampled from the student, and the student learns from that dense signal
Can we estimate the performance of the distilled student from teacher and student scale, before running OPD?
We borrow the roadmap of reward-model overoptimization
where $k_3$ is the nonnegative, unbiased estimator of token-mean reverse KL
We study Qwen2.5 base models at 0.5B, 1.5B, 3B, 7B and 14B
Vanilla-OPD minimizes reverse KL to the teacher. With the widely used zero-discount update, each sampled token gets the advantage
$$ A_t^{\mathrm{V}} = \log \pi_T(y_t \mid x, y_{\lt t}) - \log \pi_\theta(y_t \mid x, y_{\lt t}). $$Delta-OPD rewards the policy shift the teacher acquired during RL, relative to its own pre-RL checkpoint $\pi_T^{\mathrm{base}}$. A KL penalty to the student's own initialization is added. This is the common core of OPD², Direct-OPD and W2S-OPD
Off-policy distillation (OffPD) is plain SFT on teacher rollouts, $\min_\theta \mathbb{E}_{y\sim\pi_T}[-\log \pi_\theta(y\mid x)]$.
All training uses the verl framework
Every run begins with a regular useful-transfer regime, in which gold score rises approximately linearly in $d$: $G(d) = c + m\,d$. Linear fits to the first 30 checkpoints reach $R^2 \in [0.932, 0.988]$ across all 25 Vanilla-OPD runs, with slopes $m \in [0.175, 0.721]$. After the transfer endpoint $d_\mathrm{transfer}$, dynamics turn noisy and heterogeneous. Some runs keep improving more slowly, some saturate, and some regress. We define the endpoint as the last $d$ before the trajectory leaves the 95% predictive band of its initial line for three consecutive checkpoints.
Explore every trajectory below. Pick a student size, then compare how each teacher drives it.
Along a smooth training path, the gold score changes to first order in the parameter perturbation, while KL changes to second order. Taking the square root of KL puts both on the same footing:
$$ G(d) = G_0 + m\,d + O(d^2), \qquad m = \sqrt{2}\,\frac{g_0^\top h}{\sqrt{h^\top F_{\mathrm{tok}}\, h}}, $$where $g_0$ is the gold-score gradient, $h$ the update direction, and $F_{\mathrm{tok}}$ the token-weighted Fisher matrix. The slope is the Fisher-normalized alignment of the update with the gold-score gradient. How long the line persists is an empirical matter.
Takeaway. Early OPD is remarkably regular. Weak teachers generally induce smaller initial slopes for large students. Teachers closer to the student's scale raise gold score faster per unit of $d$. Where gold score regresses late, the teacher's proxy reward keeps rising, which is the signature of implicit-reward overoptimization.
In every weak-to-strong pair, the student’s peak gold score exceeds its teacher’s own score. The margin shrinks as the teacher’s size approaches the student’s. For the 7B and 14B students, the peak climbs monotonically with teacher size, from 72.7% to 81.9% and from 77.7% to 87.3%. But teacher scale helps only up to about the student’s own scale. The 0.5B student peaks at 40.8% with a 3B teacher and drops to 37.7% with a 14B teacher.
Takeaway. A compact RL expert is a cheap source of task skill for a whole model family. Same-base Vanilla-OPD peaks land within 0.5 points of direct RL at all five scales. Direct RL on the student is still the ceiling, so weak-to-strong OPD pays off by amortizing one expert, not by beating RL.
We normalize parameter counts by 1B and cap the teacher at the student’s scale, $\widetilde N_T^{\mathrm{eff}} = \min(N_T, N_S)/1\mathrm{B}$. The cap captures the saturation seen above. Parameter count alone cannot describe an undertrained teacher, so the laws also condition on the effective teacher’s remaining error $1 - G_T^{\mathrm{eff}}$:
\[1 - G_{\mathrm{peak}} = A\,\widetilde N_S^{-\alpha}\,\big(\widetilde N_T^{\mathrm{eff}}\big)^{-\beta}\,\big(1 - G_T^{\mathrm{eff}}\big)^{\zeta}, \qquad m = B\,\widetilde N_S^{-\gamma}\,\big(\widetilde N_T^{\mathrm{eff}}\big)^{\delta}\,\big(1 - G_T^{\mathrm{eff}}\big)^{-\xi}.\]| Objective | $A$ | $\alpha$ | $\beta$ | $\zeta$ | $B$ | $\gamma$ | $\delta$ | $\xi$ |
|---|---|---|---|---|---|---|---|---|
| Vanilla-OPD | 0.97 | 0.30 | −0.27 | 0.95 | 0.12 | 0.19 | −0.61 | 1.90 |
| Delta-OPD | 1.02 | 0.34 | −0.33 | 1.01 | 0.13 | 0.20 | −0.73 | 2.01 |
Two readings stand out.
Try it yourself. The presets reproduce the paper’s out-of-sample test, in which a 3B teacher checkpoint taken mid-RL is matched in score to the 1.5B RL endpoint.
The held-out checkpoint test is the sharpest check of $\beta < 0$. The 3B mid-RL teacher scores slightly higher than the 1.5B endpoint, yet its 7B student peaks lower. Only the joint law gets the order right:
| Teacher → 7B student | Teacher score | Observed peak | Scales only | Score only | Joint |
|---|---|---|---|---|---|
| 1.5B endpoint (step 580) | 63.6 | 77.5 | 77.0 | 75.6 | 77.1 |
| 3B intermediate (step 58) | 66.0 | 73.8 | 79.5 | 76.3 | 74.1 |
| 3B endpoint (step 580) | 73.6 | 80.3 | 79.5 | 78.7 | 79.6 |
How well do the laws fit overall? The next chart plots every cell’s observed value against the law’s prediction. It includes the bootstrapped-teacher cells, which are what make the teacher exponents identifiable.
| Objective | Peak law | Leave-one-scale-out RMSE ↓ | Extrapolation RMSE, student / teacher ↓ |
|---|---|---|---|
| Vanilla-OPD | Scales only | 3.43 | 0.64 / 0.75 |
| Teacher score only | 2.55 | 2.54 / 0.70 | |
| Joint | 1.66 | 0.55 / 0.68 | |
| Delta-OPD | Scales only | 2.47 | 0.21 / 0.29 |
| Teacher score only | 2.07 | 0.93 / 0.46 | |
| Joint | 0.82 | 0.20 / 0.32 |
The transfer extent does not follow a comparable law. Observed departures and censored lower bounds span $d$ of 0.20–0.36 for Vanilla-OPD and 0.27–0.34 for Delta-OPD, with medians of 0.30 and 0.29. It is better read as an approximately scale-free KL budget.
We vary one design choice at a time.
Its slope beats Vanilla-OPD in 15 of 17 shared pairs, and it reaches the higher peak in 12, mostly weak-to-strong pairs. The gain shrinks with teacher size: the 0.5B teacher adds 2–4 points, while larger teachers stay within about one point.
Pure OPD wins in every cell. An off-policy SFT cold start costs 6.6, 15.7 and 19.4 points for 3B, 7B and 14B students of the 0.5B expert, and pins them near the teacher's own score. Pure OffPD trails by up to 28.4 points.
Chaining 0.5B→1.5B→3B→7B→14B peaks below direct OPD from the 0.5B expert at every size, for example 76.8 against 77.7 at 14B. The chain's intermediate teachers score higher, yet teach worse.
The bootstrapping result reverses the gains that Burns et al. report for weak-to-strong fine-tuning
Scope. All results come from Qwen2.5 base models on math reasoning, and each curve is a single run. Parameter count also bundles data and compute, so the laws should be read at the compute-optimal settings of that model family.
If you find this work useful, please cite:
@misc{bao2026scaling,
title = {Scaling Properties of Same-Family On-Policy Distillation},
author = {Bao, Yuntai and Li, Qinfeng and Jiang, Guoqing and Chen, Liwei and
Qin, Zhiheng and Li, Xuanping and Zhang, Wenqi and Zhang, Xuhong},
year = {2026},
eprint = {2609.32722},
archiveprefix = {arXiv},
primaryclass = {cs.LG},
url = {https://arxiv.org/abs/2609.32722}
}