The multi-armed bandit problem (2012)

The multi-armed bandit problem (2012)

The other 90% of the time, we choose the lever that has the highest expectation of rewards. After many more visits, the best choice, if there is one, will have been found, and will be shown 90% of the time. In the epsilon-first strategy, you can explore 100% of the time in the beginning and once you have a good sample, switch to pure-greedy.

Source: stevehanov.ca