← Back to all articles
arXiv cs.LGOctober 2, 2026

Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models

Excerpt

arXiv:2610.00661v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states wh