arXiv cs.AIOctober 7, 2026
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
Excerpt
arXiv:2607.02781v2 Announce Type: replace-cross Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single scalar score, so explicit safety constraints must either be ignored or encoded through manually tuned penalties. We propose Lagrangian Reward Augmentation (LARA), a general inference-time alignment framework under