Almost Sure Convergence of Networked Policy Gradient over Time-Varying Networks in Markov Potential Games

Aydin, Sarper; Eksin, Ceyhun

Electrical Engineering and Systems Science > Systems and Control

arXiv:2410.20075 (eess)

[Submitted on 26 Oct 2024]

Title:Almost Sure Convergence of Networked Policy Gradient over Time-Varying Networks in Markov Potential Games

Authors:Sarper Aydin, Ceyhun Eksin

View PDF HTML (experimental)

Abstract:We propose networked policy gradient play for solving Markov potential games including continuous action and state spaces. In the decentralized algorithm, agents sample their actions from parametrized and differentiable policies that depend on the current state and other agents' policy parameters. During training, agents estimate their gradient information through two consecutive episodes, generating unbiased estimators of reward and policy score functions. Using this information, agents compute the stochastic gradients of their policy functions and update their parameters accordingly. Additionally, they update their estimates of other agents' policy parameters based on the local estimates received through a time-varying communication network. In Markov potential games, there exists a potential value function among agents with gradients corresponding to the gradients of local value functions. Using this structure, we prove the almost sure convergence of joint policy parameters to stationary points of the potential value function. We also show that the convergence rate of the networked policy gradient algorithm is $\mathcal{O}(1/\epsilon^2)$. Numerical experiments on a dynamic multi-agent newsvendor problem verify the convergence of local beliefs and gradients. It further shows that networked policy gradient play converges as fast as independent policy gradient updates, while collecting higher rewards.

Comments:	22 pages, journal version
Subjects:	Systems and Control (eess.SY); Optimization and Control (math.OC)
Cite as:	arXiv:2410.20075 [eess.SY]
	(or arXiv:2410.20075v1 [eess.SY] for this version)
	https://doi.org/10.48550/arXiv.2410.20075

Submission history

From: Sarper Aydin [view email]
[v1] Sat, 26 Oct 2024 04:36:15 UTC (5,077 KB)

Electrical Engineering and Systems Science > Systems and Control

Title:Almost Sure Convergence of Networked Policy Gradient over Time-Varying Networks in Markov Potential Games

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Systems and Control

Title:Almost Sure Convergence of Networked Policy Gradient over Time-Varying Networks in Markov Potential Games

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators