Story
arxiv_cs_lg ยท Jun 10, 2026 ยท paper
arxiv.orgJun 10, 2026
original source linked
In brief
We study policy representation learning from unlabeled multi-policy behavioral data. Each episode is generated by a fixed policy, but policy labels are unavailable. This setting appears in robotics play, demonstration...
Feed lens
agenteval