Deep Residual Reinforcement Learning

descriptionPublicationkeyboard_double_arrow_right Article , Preprint , Conference object 05 May 2020Embargo end date: 01 Jan 2019 United Kingdom Publisher:IEEE Computer SocietyJournal:International Joint Conference on Autonomous Agents and Multiagent Systems (issn: 1558-2914,

Copyright policy )Funded by:EC | CoPS

Authors: Zhang, S; Boehmer, W; Whiteson, S;

doi: 10.65109/icps4845 , 10.48550/arxiv.1905.01072

arXiv: 1905.01072

Deep Residual Reinforcement Learning

- Summary
- Subjects
- Metrics

Abstract

We revisit residual algorithms in both model-free and model-based reinforcement learning settings. We propose the bidirectional target network technique to stabilize residual algorithms, yielding a residual version of DDPG that significantly outperforms vanilla DDPG in the DeepMind Control Suite benchmark. Moreover, we find the residual algorithm an effective approach to the distribution mismatch problem in model-based planning. Compared with the existing TD(k) method, our residual-based method makes weaker assumptions about the model and yields a greater performance boost.

Country

United Kingdom

Related Organizations

Department of Computer Science University of Oxford
United Kingdom
THE CHANCELLOR, MASTERS AND SCHOLARS OF THE UNIVERSITY OF OXFORD
United Kingdom
University of Oxford
United Kingdom

Keywords

FOS: Computer and information sciences, Computer Science - Machine Learning, Artificial Intelligence (cs.AI), Computer Science - Artificial Intelligence, Statistics - Machine Learning, Machine Learning (stat.ML), Machine Learning (cs.LG)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

0

Average

Green

Fields of Science (4) View all

Fields of Science

Funded by

Related to Research communities

UArctic