Enforcing the consensus between Trajectory Optimization and Policy Learning for precise robot control

descriptionPublicationkeyboard_double_arrow_right Article , Preprint , Conference object 29 May 2023Embargo end date: 01 Jan 2022 France Publisher:IEEEJournal:2023 IEEE International Conference on Robotics and Automation (ICRA)Funded by:ANR | PRAIRIE, EC | AGIMUS

Authors: Le Lidec, Quentin; Jallet, Wilson; Laptev, Ivan; Schmid, Cordelia; Carpentier, Justin;

doi: 10.1109/icra48891.2023.10160387 , 10.48550/arxiv.2209.09006

arXiv: 2209.09006

Enforcing the consensus between Trajectory Optimization and Policy Learning for precise robot control

- Summary
- Subjects
- Metrics

Abstract

Reinforcement learning (RL) and trajectory optimization (TO) present strong complementary advantages. On one hand, RL approaches are able to learn global control policies directly from data, but generally require large sample sizes to properly converge towards feasible policies. On the other hand, TO methods are able to exploit gradient-based information extracted from simulators to quickly converge towards a locally optimal control trajectory which is only valid within the vicinity of the solution. Over the past decade, several approaches have aimed to adequately combine the two classes of methods in order to obtain the best of both worlds. Following on from this line of research, we propose several improvements on top of these approaches to learn global control policies quicker, notably by leveraging sensitivity information stemming from TO methods via Sobolev learning, and augmented Lagrangian techniques to enforce the consensus between TO and policy learning. We evaluate the benefits of these improvements on various classical tasks in robotics through comparison with existing approaches in the literature.

Country

France

Related Organizations

Centre Inria de Paris
France
French National Centre for Scientific Research
France
French Institute for Research in Computer Science and Automation
France
INSA de Toulouse
France
UNIVERSITE DE TOULOUSE
France

View all View all

Keywords

FOS: Computer and information sciences, Computer Science - Robotics, Computer Science - Machine Learning, [INFO.INFO-RB] Computer Science [cs]/Robotics [cs.RO], [INFO.INFO-LG] Computer Science [cs]/Machine Learning [cs.LG], Robotics (cs.RO), Machine Learning (cs.LG)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average