Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Other literature type . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Project proposal . 2026
License: CC BY
Data sources: Datacite
ZENODO
Project proposal . 2026
License: CC BY
Data sources: Datacite
ZENODO
Project proposal . 2026
License: CC BY
Data sources: Datacite
versions View all 3 versions
addClaim

RL-Pretrained Action Experts for Vision-Language-Action Models: Towards Physical Priors in Diffusion Policy Initialization

Authors: Scharffe;

RL-Pretrained Action Experts for Vision-Language-Action Models: Towards Physical Priors in Diffusion Policy Initialization

Abstract

Vision-Language-Action (VLA) models are promising as they allow to interact with a robot through natural language and are very generalist policies. VLAs attach a randomly-initialized Diffusion Transformer (DiT) to a pretrained Vision-Language Model (VLM) and train it via supervised fine-tuning (SFT) on real or synthetic demonstration data. In parallel, the development of modern physics simulator (MuJoCo, Isaac Lab) combined with reinforcement learning (RL) has produced highly capable control policies especially for the locomotion of very complex robots like humanoids or dexterous manipulation. One of the main bottleneck of physical AI is the lack of data. The main advantage of RL training is that it can be run inside a simulator. I propose to use RL to pretrain a VLA in a simulator to let it build a physical knowledge a priori. Unlike recent work that RL-pretrains the action expert on a single downstream manipulation task, the physical prior I target is task-agnostic: it is trained once against a generic physics reward and transfers across tasks and embodiments, including classical whole-body skills such as walking. To do so, it requires to initialize the DiT. I show that this is feasible with two established architectures: a causal transformer trained by standard PPO, or a diffusion policy trained by DPPO or DDiffPG.

Keywords

Machine Learning, Sensors, Control engineering, Reinforcement learning, Computer vision, Robotics, Natural Language Processing

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green