Story

arxiv_cs_lg ยท Apr 20, 2026 ยท paper

Source brief

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

arxiv.orgApr 20, 2026
original source linked

In brief

Uniform Discrete Diffusion Model (UDM) has recently emerged as a promising paradigm for discrete generative modeling; however, its integration with reinforcement learning remains largely unexplored. We observe that na...

Continue reading

Read the original at arxiv.org โ†’Open in live feed