Learning Goal-Conditioned Policies Offline with Self-Supervised Reward Shaping

Mezghani, Lina; Sukhbaatar, Sainbayar; Bojanowski, Piotr; Lazaric, Alessandro; Alahari, Karteek

Computer Science > Robotics

arXiv:2301.02099 (cs)

[Submitted on 5 Jan 2023]

Title:Learning Goal-Conditioned Policies Offline with Self-Supervised Reward Shaping

Authors:Lina Mezghani, Sainbayar Sukhbaatar, Piotr Bojanowski, Alessandro Lazaric, Karteek Alahari

View PDF

Abstract:Developing agents that can execute multiple skills by learning from pre-collected datasets is an important problem in robotics, where online interaction with the environment is extremely time-consuming. Moreover, manually designing reward functions for every single desired skill is prohibitive. Prior works targeted these challenges by learning goal-conditioned policies from offline datasets without manually specified rewards, through hindsight relabelling. These methods suffer from the issue of sparsity of rewards, and fail at long-horizon tasks. In this work, we propose a novel self-supervised learning phase on the pre-collected dataset to understand the structure and the dynamics of the model, and shape a dense reward function for learning policies offline. We evaluate our method on three continuous control tasks, and show that our model significantly outperforms existing approaches, especially on tasks that involve long-term planning.

Comments:	Code: this https URL
Subjects:	Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2301.02099 [cs.RO]
	(or arXiv:2301.02099v1 [cs.RO] for this version)
	https://doi.org/10.48550/arXiv.2301.02099
Journal reference:	6th Conference on Robot Learning (CoRL 2022)

Submission history

From: Lina Mezghani [view email]
[v1] Thu, 5 Jan 2023 15:07:10 UTC (2,904 KB)

Computer Science > Robotics

Title:Learning Goal-Conditioned Policies Offline with Self-Supervised Reward Shaping

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Robotics

Title:Learning Goal-Conditioned Policies Offline with Self-Supervised Reward Shaping

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators