← Back to all articles
arXiv cs.LGOctober 2, 2026

Bellman-Certified Rounding for Sparse Policy Deployment in MDPs

Excerpt

arXiv:2610.00325v1 Announce Type: new Abstract: Continuous policy optimization may spread an update across many states, even when deployment permits only a few complete state-level changes. We study how much discounted return can be retained when continuous row mixtures are rounded to sparse binary policies in finite MDPs. Policy-dependent visitation couples the row edits, while long horizons make global curvature bounds conservative. From $2d+2$ Bellman solves, we derive reusable envelopes that