Powered by OpenAIRE graph
Found an issue? Give us feedback
addClaim

Data for Arnold: A multi-task, multi-embodiment muscle transformer policy

Authors: Chiappa, Alberto Silvio; An, Boshi; Simos, Merkourios; Li, Chengkun; Mathis, Alexander;

Data for Arnold: A multi-task, multi-embodiment muscle transformer policy

Abstract

Supplementary Data — Trained Model Checkpoints and Benchmark Results==================================================================== This archive set accompanies an anonymous manuscript submission. It containsthe final trained model weights and the corresponding benchmark evaluationresults for every model reported in the paper. Contents--------1. arnold-final-checkpoints.tar.gz Trained model weights. Unpacks to: final_checkpoints/ ~2.8 GB. One subdirectory per model (see "Models" below), each containing per-seed subdirectories (seed_0, seed_1, ...). A typical seed directory contains: - rl_model__steps.zip Stable-Baselines3 policy checkpoint - rl_model_vecnormalize__steps.pkl VecNormalize statistics - model_config.json / args.json training + model configuration - _config.json per-task environment configs - vocabulary.json sensorimotor vocabulary (where used) - main_bc_ppo_multi_task.py training entry point (snapshot) - PPO_0/ training logs / tensorboard events 2. arnold-benchmark-results.tar.gz Benchmark evaluation results. Unpacks to: benchmarks/final_benchmarks/ ~4 MB. Same model/seed layout as above; each seed directory holds: - seed___results.json per-task benchmark scores - *.log evaluation run logs 3. arnold-benchmark-results-extra.tar.gz Benchmark evaluation results. Unpacks to: benchmarks/final_benchmarks_extra/ ~11 MB. Contains files used by plotting or other analysis in the paper. Model Description arnold Proposed method (multi-task compositional policy) obc On-policy behavior cloning baseline obc_10_tasks OBC trained on the 10-task subset obc_m / obc_s / obc_xs OBC capacity variants (medium / small / extra-small) obc_st OBC single-task specialists obc_task_sv OBC, task-specific sensorimotor vocabulary variant obc_ppo OBC with PPO fine-tuning obc_wo_obs_norm OBC ablation without observation normalization bc Behavior-cloning baseline ppo_t Transformer-based PPO (no sensorimotor vocabulary) ppo_t_sv Transformer-based PPO with sensorimotor vocabulary mt-ppo Multi-task PPO mt-sac Multi-task SAC

Keywords

Reinforcement learning, motor control, skill learning

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average