arXiv cs.LGOctober 7, 2026
Uncovering Cross-Objective Interference in Multi-Objective Alignment
Excerpt
arXiv:2602.06869v3 Announce Type: replace-cross Abstract: We study a persistent failure mode in multi-objective alignment for large language models (LLMs), in which scalarized training improves only some objectives while the others degrade. We formalize this phenomenon as cross-objective interference and, to our knowledge, conduct the first systematic study of scalarization algorithms for multi-objective LLM alignment. The study shows that interference is pervasive across algorithms yet strongly