Scaling Properties of Same-Family On-Policy Distillation

how much RL-acquired capability transfers across model scales via on-policy distillation, how fast, and how to predict it before training.

Overview: OPD setups, two-phase training dynamics, and the joint power law
Overview. Left: the three teacher–student setups, with post-RL experts as teachers and SFT bases as students. Center: gold score rises linearly with KL-indexed training progress $d$ during the useful-transfer phase, after which dynamics turn noisy. Right: the fitted joint power law predicts peak gold error from teacher and student scale.
TL;DR

On-policy distillation (OPD) from a small RL-trained expert reliably lifts much larger students, and the outcome is predictable before training: gold score first rises linearly in $\sqrt{\mathrm{KL}}$, and the peak follows a power law in student size, teacher size, and teacher score.

0.93–0.99
$R^2$ of a straight line $G(d)=c+md$ over the first 30 checkpoints, across all 25 runs
10 / 10
weak-to-strong pairs whose student peaks above its teacher's own score
≤ 0.7 pt
error of the joint peak law when extrapolating to the held-out largest student or teacher
−19.4 pt
peak cost of one epoch of off-policy SFT cold start (0.5B teacher → 14B student)

The question

Reinforcement learning (RL) instills strong reasoning skills in LLMs, but it is expensive to repeat at every model size. On-policy distillation offers a shortcut: a teacher scores every token of rollouts sampled from the student, and the student learns from that dense signal. We ask one central question:

Can we estimate the performance of the distilled student from teacher and student scale, before running OPD?

We borrow the roadmap of reward-model overoptimization. OPD also optimizes the student against a proxy, namely the token-level implicit reward of a fixed teacher. So we track held-out accuracy, the gold score $G$, as a function of how far the student has moved from its initialization:

\[\begin{aligned} d &\mathrel{:=} \sqrt{k_3(\pi_\theta, \pi_\mathrm{ref})}, \qquad k_3 = \mathbb{E}\Big[\tfrac{1}{|y|}\textstyle\sum_t e^{\delta_t} - \delta_t - 1\Big], \\[6pt] \delta_t &= \log \pi_\mathrm{ref}(y_t \mid x, y_{\lt t}) - \log \pi_\theta(y_t \mid x, y_{\lt t}), \end{aligned}\]

where $k_3$ is the nonnegative, unbiased estimator of token-mean reverse KL.

Setup

We study Qwen2.5 base models at 0.5B, 1.5B, 3B, 7B and 14B. Every model first gets a short SFT phase. Teachers are then trained with GRPO on the mixed GSM8K + MATH training split (14.8K problems), and gold score is accuracy on the mixed held-out test split (6.3K problems). Crossing all five teachers with all five students gives 25 teacher–student pairs. The checkpoints from these runs, including teachers, students and baselines, are released on Hugging Face:

↗ Weak-to-strong
small RL expert teaches a larger SFT student (10 pairs)
→ Same-base
teacher and student share a base size (5 pairs)
↘ Strong-to-weak
classic distillation into a smaller student (10 pairs)
The objectives we compare (click to expand)

Vanilla-OPD minimizes reverse KL to the teacher. With the widely used zero-discount update, each sampled token gets the advantage

$$ A_t^{\mathrm{V}} = \log \pi_T(y_t \mid x, y_{\lt t}) - \log \pi_\theta(y_t \mid x, y_{\lt t}). $$

Delta-OPD rewards the policy shift the teacher acquired during RL, relative to its own pre-RL checkpoint $\pi_T^{\mathrm{base}}$. A KL penalty to the student's own initialization is added. This is the common core of OPD², Direct-OPD and W2S-OPD:

$$ A_t^{\Delta} = \log \pi_T(y_t \mid x, y_{\lt t}) - \log \pi_T^{\mathrm{base}}(y_t \mid x, y_{\lt t}). $$

Off-policy distillation (OffPD) is plain SFT on teacher rollouts, $\min_\theta \mathbb{E}_{y\sim\pi_T}[-\log \pi_\theta(y\mid x)]$.

All training uses the verl framework. Each OPD run lasts at most 580 updates, which is ten epochs, and the held-out set is evaluated periodically.

Finding 1

Transfer starts with a straight line in √KL

Every run begins with a regular useful-transfer regime, in which gold score rises approximately linearly in $d$: $G(d) = c + m\,d$. Linear fits to the first 30 checkpoints reach $R^2 \in [0.932, 0.988]$ across all 25 Vanilla-OPD runs, with slopes $m \in [0.175, 0.721]$. After the transfer endpoint $d_\mathrm{transfer}$, dynamics turn noisy and heterogeneous. Some runs keep improving more slowly, some saturate, and some regress. We define the endpoint as the last $d$ before the trajectory leaves the 95% predictive band of its initial line for three consecutive checkpoints.

Explore every trajectory below. Pick a student size, then compare how each teacher drives it.

Student
Objective
Interactive: gold-score trajectories. Each colored curve is one OPD run, with darker colors for larger teachers. Dashed lines are the linear fits to the first 30 checkpoints (40 for Delta-OPD), drawn up to each run's transfer endpoint. A vertical tick marks runs that depart from their line within the observed range. Stars mark peaks, and the dotted level is direct RL on the student itself. Dots replay training: all runs advance one update at a time, so runs that cover more $d$ per update move faster along the axis. Every curve is one run.
Why does √KL linearize transfer?

Along a smooth training path, the gold score changes to first order in the parameter perturbation, while KL changes to second order. Taking the square root of KL puts both on the same footing:

$$ G(d) = G_0 + m\,d + O(d^2), \qquad m = \sqrt{2}\,\frac{g_0^\top h}{\sqrt{h^\top F_{\mathrm{tok}}\, h}}, $$

where $g_0$ is the gold-score gradient, $h$ the update direction, and $F_{\mathrm{tok}}$ the token-weighted Fisher matrix. The slope is the Fisher-normalized alignment of the update with the gold-score gradient. How long the line persists is an empirical matter.

Takeaway. Early OPD is remarkably regular. Weak teachers generally induce smaller initial slopes for large students. Teachers closer to the student's scale raise gold score faster per unit of $d$. Where gold score regresses late, the teacher's proxy reward keeps rising, which is the signature of implicit-reward overoptimization.

Finding 2

Small experts lift much larger students

In every weak-to-strong pair, the student’s peak gold score exceeds its teacher’s own score. The margin shrinks as the teacher’s size approaches the student’s. For the 7B and 14B students, the peak climbs monotonically with teacher size, from 72.7% to 81.9% and from 77.7% to 87.3%. But teacher scale helps only up to about the student’s own scale. The 0.5B student peaks at 40.8% with a 3B teacher and drops to 37.7% with a 14B teacher.

Objective
Interactive: peak gold score (%) for every teacher–student pair. Cells left of the diagonal are weak-to-strong, the diagonal is same-base, and cells to the right are strong-to-weak. Hover over a cell to compare the peak with the teacher's own score and with direct RL on the student. Delta-OPD covers 17 of the 25 pairs.

Takeaway. A compact RL expert is a cheap source of task skill for a whole model family. Same-base Vanilla-OPD peaks land within 0.5 points of direct RL at all five scales. Direct RL on the student is still the ceiling, so weak-to-strong OPD pays off by amortizing one expert, not by beating RL.

Finding 3

Power laws predict the peak before training

We normalize parameter counts by 1B and cap the teacher at the student’s scale, $\widetilde N_T^{\mathrm{eff}} = \min(N_T, N_S)/1\mathrm{B}$. The cap captures the saturation seen above. Parameter count alone cannot describe an undertrained teacher, so the laws also condition on the effective teacher’s remaining error $1 - G_T^{\mathrm{eff}}$:

\[1 - G_{\mathrm{peak}} = A\,\widetilde N_S^{-\alpha}\,\big(\widetilde N_T^{\mathrm{eff}}\big)^{-\beta}\,\big(1 - G_T^{\mathrm{eff}}\big)^{\zeta}, \qquad m = B\,\widetilde N_S^{-\gamma}\,\big(\widetilde N_T^{\mathrm{eff}}\big)^{\delta}\,\big(1 - G_T^{\mathrm{eff}}\big)^{-\xi}.\]
Objective $A$ $\alpha$ $\beta$ $\zeta$ $B$ $\gamma$ $\delta$ $\xi$
Vanilla-OPD 0.97 0.30 −0.27 0.95 0.12 0.19 −0.61 1.90
Delta-OPD 1.02 0.34 −0.33 1.01 0.13 0.20 −0.73 2.01
Fitted joint power-law coefficients. The left four columns are for peak capability and the right four for useful-transfer rate. On the RL-endpoint grid, teacher size and teacher score are almost perfectly collinear. Teachers from the bootstrapped chains, which score below the size trend, break that collinearity and identify every exponent.

Two readings stand out.

Try it yourself. The presets reproduce the paper’s out-of-sample test, in which a 3B teacher checkpoint taken mid-RL is matched in score to the 1.5B RL endpoint.

Vanilla-OPD peak
slope m ≈
Delta-OPD peak
slope m ≈
Interactive: the joint peak law. Predicted peak gold score and initial slope from Equation 6 of the paper, using the full-precision fitted coefficients. The curve sweeps student size for the chosen teacher. The fits cover 0.5B to 14B Qwen2.5 models on math, so values outside that range are extrapolations.

The held-out checkpoint test is the sharpest check of $\beta < 0$. The 3B mid-RL teacher scores slightly higher than the 1.5B endpoint, yet its 7B student peaks lower. Only the joint law gets the order right:

Teacher → 7B student Teacher score Observed peak Scales only Score only Joint
1.5B endpoint (step 580) 63.6 77.5 77.0 75.6 77.1
3B intermediate (step 58) 66.0 73.8 79.5 76.3 74.1
3B endpoint (step 580) 73.6 80.3 79.5 78.7 79.6
7B student trajectories under the 1.5B endpoint, 3B intermediate and 3B endpoint teachers
Intermediate-teacher validation. Gold score against $d$ for the 7B student distilled from the 1.5B RL endpoint, the score-matched 3B checkpoint at step 58, and the 3B RL endpoint. Stars mark peaks, and dashed levels mark each teacher's own gold score.

How well do the laws fit overall? The next chart plots every cell’s observed value against the law’s prediction. It includes the bootstrapped-teacher cells, which are what make the teacher exponents identifiable.

Target
Interactive: observed against predicted. Predictions come from the full-fit joint laws. The dotted diagonal is a perfect prediction, and open diamonds mark cells whose teacher is itself an OPD product of a bootstrapped chain. The rate fits are noisier, with log-space $R^2$ of 0.59 for Vanilla-OPD and 0.76 for Delta-OPD. Read the rate law as an interpretable summary rather than a precise predictor.
Objective Peak law Leave-one-scale-out RMSE ↓ Extrapolation RMSE, student / teacher ↓
Vanilla-OPD Scales only 3.43 0.64 / 0.75
  Teacher score only 2.55 2.54 / 0.70
  Joint 1.66 0.55 / 0.68
Delta-OPD Scales only 2.47 0.21 / 0.29
  Teacher score only 2.07 0.93 / 0.46
  Joint 0.82 0.20 / 0.32
Validation of the peak laws, in accuracy points. Leave-one-scale-out RMSE refits each law with every cell sharing one student or teacher scale withheld. Extrapolation RMSE withholds the largest student or teacher scale.

The transfer extent does not follow a comparable law. Observed departures and censored lower bounds span $d$ of 0.20–0.36 for Vanilla-OPD and 0.27–0.34 for Delta-OPD, with medians of 0.30 and 0.29. It is better read as an approximately scale-free KL budget.

Finding 4

Design choices: what helps and what hurts

We vary one design choice at a time.

Peak gold score for Delta-OPD, Vanilla-OPD and direct RL

Delta-OPD transfers faster ✓

Its slope beats Vanilla-OPD in 15 of 17 shared pairs, and it reaches the higher peak in 12, mostly weak-to-strong pairs. The gain shrinks with teacher size: the 0.5B teacher adds 2–4 points, while larger teachers stay within about one point.

Peak gold score under three degrees of on-policy supervision

Stay on-policy ✓

Pure OPD wins in every cell. An off-policy SFT cold start costs 6.6, 15.7 and 19.4 points for 3B, 7B and 14B students of the 0.5B expert, and pins them near the teacher's own score. Pure OffPD trails by up to 28.4 points.

Bootstrapped 0.5B to 14B chains against direct OPD

Don't bootstrap ✗

Chaining 0.5B→1.5B→3B→7B→14B peaks below direct OPD from the 0.5B expert at every size, for example 76.8 against 77.7 at 14B. The chain's intermediate teachers score higher, yet teach worse.

The bootstrapping result reverses the gains that Burns et al. report for weak-to-strong fine-tuning. It echoes recent OPD findings that a higher-scoring teacher helps only when it offers new capabilities. A teacher’s score and its teaching value can come apart. For example, the 1.5B OPD product scores 53.0% against the 0.5B RL expert’s 39.8%, yet its 3B student peaks lower.

Takeaways for practitioners

Scope. All results come from Qwen2.5 base models on math reasoning, and each curve is a single run. Parameter count also bundles data and compute, so the laws should be read at the compute-optimal settings of that model family.

Citation

If you find this work useful, please cite:

@misc{bao2026scaling,
  title         = {Scaling Properties of Same-Family On-Policy Distillation},
  author        = {Bao, Yuntai and Li, Qinfeng and Jiang, Guoqing and Chen, Liwei and
                   Qin, Zhiheng and Li, Xuanping and Zhang, Wenqi and Zhang, Xuhong},
  year          = {2026},
  eprint        = {2609.32722},
  archiveprefix = {arXiv},
  primaryclass  = {cs.LG},
  url           = {https://arxiv.org/abs/2609.32722}
}