A Simple and Optimal Policy Design with Safety against Heavy-tailed Risk for Stochastic Bandits

Simchi-Levi, David; Zheng, Zeyu; Zhu, Feng

Statistics > Machine Learning

arXiv:2206.02969v5 (stat)

[Submitted on 7 Jun 2022 (v1), revised 15 Nov 2022 (this version, v5), latest version 22 Jul 2024 (v6)]

Title:A Simple and Optimal Policy Design with Safety against Heavy-tailed Risk for Stochastic Bandits

Authors:David Simchi-Levi, Zeyu Zheng, Feng Zhu

View PDF

Abstract:We study the stochastic multi-armed bandit problem and design new policies that enjoy both worst-case optimality for expected regret and light-tailed risk for regret distribution. Starting from the two-armed bandit setting with time horizon $T$, we propose a simple policy and prove that the policy (i) enjoys the worst-case optimality for the expected regret at order $O(\sqrt{T\ln T})$ and (ii) has the worst-case tail probability of incurring a linear regret decay at an exponential rate $\exp(-\Omega(\sqrt{T}))$, a rate that we prove to be best achievable for all worst-case optimal policies. Briefly, our proposed policy achieves a delicate balance between doing more exploration at the beginning of the time horizon and doing more exploitation when approaching the end, compared to the standard Successive Elimination policy and Upper Confidence Bound policy. We then improve the policy design and analysis to work for the general $K$-armed bandit setting. Specifically, the worst-case probability of incurring a regret larger than any $x>0$ is upper bounded by $\exp(-\Omega(x/\sqrt{KT}))$. We then enhance the policy design to accommodate the "any-time" setting where $T$ is not known a priori, and prove equivalently desired policy performances as compared to the "fixed-time" setting with known $T$. A brief account of numerical experiments is conducted to illustrate the theoretical findings. We conclude by extending our proposed policy design to the general stochastic linear bandit setting and proving that the policy leads to both worst-case optimality in terms of expected regret order and light-tailed risk on the regret distribution.

Comments:	Preliminary version appeared in NeurIPS 2022
Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
Cite as:	arXiv:2206.02969 [stat.ML]
	(or arXiv:2206.02969v5 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2206.02969

Submission history

From: Feng Zhu [view email]
[v1] Tue, 7 Jun 2022 02:10:30 UTC (1,654 KB)
[v2] Fri, 10 Jun 2022 02:16:45 UTC (1,659 KB)
[v3] Mon, 4 Jul 2022 16:15:22 UTC (1,756 KB)
[v4] Tue, 1 Nov 2022 18:43:11 UTC (2,418 KB)
[v5] Tue, 15 Nov 2022 02:43:45 UTC (2,414 KB)
[v6] Mon, 22 Jul 2024 14:45:09 UTC (3,414 KB)

Statistics > Machine Learning

Title:A Simple and Optimal Policy Design with Safety against Heavy-tailed Risk for Stochastic Bandits

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:A Simple and Optimal Policy Design with Safety against Heavy-tailed Risk for Stochastic Bandits

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators