-
Scaling Properties of Same-Family On-Policy Distillation
how much RL-acquired capability transfers across model scales via on-policy distillation, how fast, and how to predict it before training.
-
Things I Wish Someone Had Told Me before My PhD
a halfway reflection on my PhD, sharing my personal experiences (not lecture).
-
On-Policy Distillation
an informal review of on-policy distillation.
-
Rubric-based Rewards in Reinforcement Learning
an informal review of RL with rubrics as rewards.
-
Training Prompt-only Steering Vectors in a Principled Manner
our recent work on prompt-only SV and SV training dynamics.