Contextual Bandit Algorithms with Supervised Learning Guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, Robert E. Schapire
Abstract
We address the problem of competing with any large set of $N$ policies in the non-stochastic bandit setting, where the learner must repeatedly select among $K$ actions but observes only the reward of the chosen action.
原文 arXiv:1002.4058;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1002.4058v3