← Back to all articles
arXiv cs.LGOctober 7, 2026

Reinforcement Learning over Predictive Distributions for LLM Regression

Excerpt

arXiv:2605.20740v2 Announce Type: replace Abstract: Large language models (LLMs) have emerged as flexible regressors capable of predicting real-valued quantities from heterogeneous inputs. Yet most LLM regression objectives optimize predictions independently, often yielding poor calibration. We introduce Distribution-Aware Reward (DAR), an on-policy reinforcement learning objective that instead jointly evaluates the empirical predictive distribution formed by multiple predictions for the same in