← Back to all articles
arXiv cs.LGOctober 2, 2026

dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models

Excerpt

arXiv:2605.09291v2 Announce Type: replace Abstract: Discrete flow models (DFMs) are a class of flexible generative models for generating discrete data, and diffusion large language models (dLLMs) can be viewed as a special case with a specific choice of a mixture path and a masked source distribution. While several recent works have explored reinforcement learning for dLLMs, its application to more general discrete flow models remains underexplored. In this work, we present discrete Flow-GRPO (d