← Back to all articles
arXiv cs.CLOctober 7, 2026

Unbiased Reward Modeling from Implicit Feedback for LLM Alignment

Excerpt

arXiv:2603.23184v2 Announce Type: replace Abstract: Despite the success of reinforcement learning from human feedback (RLHF), existing reward modeling methods largely rely on explicit feedback, which is costly to collect and difficult to scale. This work studies implicit reward modeling, learning reward models from implicit user feedback, such as clicks, copies and skips. While scalable and cost-effective, implicit feedback poses two key challenges: It lacks definitive negative samples, which ma