Composing Entropic Policies using Divergence Correction

Hunt, Jonathan J; Barreto, Andre; Lillicrap, Timothy P; Heess, Nicolas

Computer Science > Machine Learning

arXiv:1812.02216 (cs)

[Submitted on 5 Dec 2018 (v1), last revised 5 Jul 2019 (this version, v2)]

Title:Composing Entropic Policies using Divergence Correction

Authors:Jonathan J Hunt, Andre Barreto, Timothy P Lillicrap, Nicolas Heess

View PDF

Abstract:Composing previously mastered skills to solve novel tasks promises dramatic improvements in the data efficiency of reinforcement learning. Here, we analyze two recent works composing behaviors represented in the form of action-value functions and show that they perform poorly in some situations. As part of this analysis, we extend an important generalization of policy improvement to the maximum entropy framework and introduce an algorithm for the practical implementation of successor features in continuous action spaces. Then we propose a novel approach which addresses the failure cases of prior work and, in principle, recovers the optimal policy during transfer. This method works by explicitly learning the (discounted, future) divergence between base policies. We study this approach in the tabular case and on non-trivial continuous control problems with compositional structure and show that it outperforms or matches existing methods across all tasks considered.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1812.02216 [cs.LG]
	(or arXiv:1812.02216v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1812.02216

Submission history

From: Jonathan Hunt [view email]
[v1] Wed, 5 Dec 2018 20:43:29 UTC (826 KB)
[v2] Fri, 5 Jul 2019 16:22:52 UTC (2,634 KB)

Computer Science > Machine Learning

Title:Composing Entropic Policies using Divergence Correction

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Composing Entropic Policies using Divergence Correction

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators