Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
We show that on-policy distillation improves largely by suppressing low-probability student tokens rather than transferring teacher knowledge, and introduce OPSA, a supervision-free method that turns token uncertainty into self-improvement.
. Previously, I worked as a research assistant in the MLDM Lab's Multimodal Vision Processing (MVP) Group, under the guidance of
