arXiv cs.LGAugust 18, 2026
Sequential Batch Learning in Finite-Action Linear Contextual Bandits
Excerpt
arXiv:2004.06321v2 Announce Type: replace Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can only observe outcomes for the individuals within a batch at the batch's end. Compared with both standard online contextual-bandit learning and offline policy learning in contextual bandits, this sequential batch learning problem