← Back to all articles
arXiv cs.AIAugust 18, 2026

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

Excerpt

arXiv:2608.16831v1 Announce Type: new Abstract: Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evaluate-and-improve structure of generalized policy iteration. PIHF uses a pretrained language model as its execution substrate and move