Skip to content
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation · ZeroHour