arXiv cs.CLOctober 7, 2026
Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
Excerpt
arXiv:2609.33155v2 Announce Type: replace Abstract: Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes acro