Story

arxiv_cs_ai ยท Aug 19, 2026 ยท paper

Source brief

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

arxiv.orgAug 19, 2026
original source linked

In brief

On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible resp...

Feed lens
eval

Continue reading

Read the original at arxiv.org โ†’Open in live feedRead that dayโ€™s brief

Earlier in this thread 4 items