Story
arxiv_cs_cl ยท Aug 17, 2026 ยท paper
arxiv.orgAug 17, 2026
original source linked
In brief
Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not alwa...
Feed lens
evaluation