Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

Zhang, Kaiqing; Yang, Zhuoran; Başar, Tamer

Computer Science > Machine Learning

arXiv:1906.00729 (cs)

[Submitted on 31 May 2019 (v1), last revised 10 Feb 2021 (this version, v4)]

Title:Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

Authors:Kaiqing Zhang, Zhuoran Yang, Tamer Başar

View PDF

Abstract:We study the global convergence of policy optimization for finding the Nash equilibria (NE) in zero-sum linear quadratic (LQ) games. To this end, we first investigate the landscape of LQ games, viewing it as a nonconvex-nonconcave saddle-point problem in the policy space. Specifically, we show that despite its nonconvexity and nonconcavity, zero-sum LQ games have the property that the stationary point of the objective function with respect to the linear feedback control policies constitutes the NE of the game. Building upon this, we develop three projected nested-gradient methods that are guaranteed to converge to the NE of the game. Moreover, we show that all of these algorithms enjoy both globally sublinear and locally linear convergence rates. Simulation results are also provided to illustrate the satisfactory convergence properties of the algorithms. To the best of our knowledge, this work appears to be the first one to investigate the optimization landscape of LQ games, and provably show the convergence of policy optimization methods to the Nash equilibria. Our work serves as an initial step toward understanding the theoretical aspects of policy-based reinforcement learning algorithms for zero-sum Markov games in general.

Comments:	Fixed some typos, addressed some comments from NeurIPS reviews
Subjects:	Machine Learning (cs.LG); Computer Science and Game Theory (cs.GT); Systems and Control (eess.SY); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as:	arXiv:1906.00729 [cs.LG]
	(or arXiv:1906.00729v4 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1906.00729

Submission history

From: Kaiqing Zhang [view email]
[v1] Fri, 31 May 2019 17:28:25 UTC (802 KB)
[v2] Mon, 28 Oct 2019 11:05:45 UTC (1,430 KB)
[v3] Sun, 14 Jun 2020 06:36:33 UTC (1,442 KB)
[v4] Wed, 10 Feb 2021 21:30:42 UTC (2,887 KB)

Computer Science > Machine Learning

Title:Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Policy Optimization Provably Converges to Nash Equilibria in Zero-Sum Linear Quadratic Games

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators