← Back to all articles
arXiv cs.LGAugust 18, 2026

Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates

Excerpt

arXiv:2606.00984v2 Announce Type: replace-cross Abstract: We study linear contextual bandits under rare parameter updates: the learner may incorporate reward feedback into its parameter estimate only at a small number of update times, while still observing contexts online and selecting actions sequentially. This viewpoint clarifies a practical distinction that is often blurred in the literature: many "strictly batched" methods additionally restrict within-interval context adaptivity, meaning tha