arXiv cs.LGOctober 7, 2026
Axiom Satisfiability of Linear Rewards in Alignment
Excerpt
arXiv:2610.06892v1 Announce Type: cross Abstract: Learning from human preference data is the dominant route to aligning language models with human values. In linear social choice, where rewards are linear in a fixed feature representation of prompt-response pairs, Ge et al.[2024] show that fitting such a reward by minimizing any non-decreasing convex loss, including BTL, fails PO and PMC. Moreover, no rule that reads only the majority relation can satisfy PO once the output is required to be lin