Distributional Off-policy Evaluation with Bellman Residual Minimization

Hong, Sungee; Qi, Zhengling; Wong, Raymond K. W.

Statistics > Machine Learning

arXiv:2402.01900v1 (stat)

[Submitted on 2 Feb 2024 (this version), latest version 12 Mar 2025 (v3)]

Title:Distributional Off-policy Evaluation with Bellman Residual Minimization

Authors:Sungee Hong, Zhengling Qi, Raymond K. W. Wong

View PDF

Abstract:We consider the problem of distributional off-policy evaluation which serves as the foundation of many distributional reinforcement learning (DRL) algorithms. In contrast to most existing works (that rely on supremum-extended statistical distances such as supremum-Wasserstein distance), we study the expectation-extended statistical distance for quantifying the distributional Bellman residuals and show that it can upper bound the expected error of estimating the return distribution. Based on this appealing property, by extending the framework of Bellman residual minimization to DRL, we propose a method called Energy Bellman Residual Minimizer (EBRM) to estimate the return distribution. We establish a finite-sample error bound for the EBRM estimator under the realizability assumption. Furthermore, we introduce a variant of our method based on a multi-step bootstrapping procedure to enable multi-step extension. By selecting an appropriate step level, we obtain a better error bound for this variant of EBRM compared to a single-step EBRM, under some non-realizability settings. Finally, we demonstrate the superior performance of our method through simulation studies, comparing with several existing methods.

Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as:	arXiv:2402.01900 [stat.ML]
	(or arXiv:2402.01900v1 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2402.01900

Submission history

From: Sungee Hong [view email]
[v1] Fri, 2 Feb 2024 20:59:29 UTC (636 KB)
[v2] Thu, 17 Oct 2024 03:26:07 UTC (952 KB)
[v3] Wed, 12 Mar 2025 04:52:57 UTC (2,054 KB)

Statistics > Machine Learning

Title:Distributional Off-policy Evaluation with Bellman Residual Minimization

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Distributional Off-policy Evaluation with Bellman Residual Minimization

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators