Making Deep Q-learning methods robust to time discretization

Tallec, Corentin; Blier, Léonard; Ollivier, Yann

Computer Science > Machine Learning

arXiv:1901.09732 (cs)

[Submitted on 28 Jan 2019 (v1), last revised 29 Jan 2019 (this version, v2)]

Title:Making Deep Q-learning methods robust to time discretization

Authors:Corentin Tallec, Léonard Blier, Yann Ollivier

View PDF

Abstract:Despite remarkable successes, Deep Reinforcement Learning (DRL) is not robust to hyperparameterization, implementation details, or small environment changes (Henderson et al. 2017, Zhang et al. 2018). Overcoming such sensitivity is key to making DRL applicable to real world problems. In this paper, we identify sensitivity to time discretization in near continuous-time environments as a critical factor; this covers, e.g., changing the number of frames per second, or the action frequency of the controller. Empirically, we find that Q-learning-based approaches such as Deep Q- learning (Mnih et al., 2015) and Deep Deterministic Policy Gradient (Lillicrap et al., 2015) collapse with small time steps. Formally, we prove that Q-learning does not exist in continuous time. We detail a principled way to build an off-policy RL algorithm that yields similar performances over a wide range of time discretizations, and confirm this robustness empirically.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1901.09732 [cs.LG]
	(or arXiv:1901.09732v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1901.09732

Submission history

From: Léonard Blier [view email]
[v1] Mon, 28 Jan 2019 15:25:10 UTC (10,314 KB)
[v2] Tue, 29 Jan 2019 15:10:27 UTC (5,154 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2019-01

Change to browse by:

cs
stat
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Corentin Tallec
Léonard Blier
Yann Ollivier

export BibTeX citation

Computer Science > Machine Learning

Title:Making Deep Q-learning methods robust to time discretization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Making Deep Q-learning methods robust to time discretization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators