Story
arxiv_cs_ai ยท Aug 19, 2026 ยท paper
Source brief
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
arxiv.orgAug 19, 2026
original source linked
In brief
On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible resp...
Feed lens
eval