Story

arxiv_cs_lg ยท May 11, 2026 ยท paper

Source brief

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

arxiv.orgMay 11, 2026
original source linked

In brief

Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyond the regime of verifiable rewards. Their outputs lack ground-tr...

Continue reading

Read the original at arxiv.org โ†’Open in live feed

Earlier in this thread 4 items