arXiv cs.LGOctober 7, 2026
Reinforcement Learning over Predictive Distributions for LLM Regression
Excerpt
arXiv:2605.20740v2 Announce Type: replace Abstract: Large language models (LLMs) have emerged as flexible regressors capable of predicting real-valued quantities from heterogeneous inputs. Yet most LLM regression objectives optimize predictions independently, often yielding poor calibration. We introduce Distribution-Aware Reward (DAR), an on-policy reinforcement learning objective that instead jointly evaluates the empirical predictive distribution formed by multiple predictions for the same in