A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation
Not provided in the content
Abstract
The paper demonstrates that using a shared learning rate in selective on-policy distillation is not a neutral control, affecting the performance of different selectors significantly.
Reality Card
The study reveals that the shared learning rate influences the performance of selective on-policy distillation, with significant differences in outcomes based on the learning rate used.
The dense-versus-selective verdict shows a 10.1 pp difference at lr=1e-4 and a 5.1 pp difference at 5e-5, indicating a 2.0x variation based on the learning rate.
The results may not reproduce under different conditions, as indicated by the varying performance on MATH-500 and the specific context of LoRA.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.