
This release expands FluxVLA's real-robot deployment support with single- and dual-arm Franka and Oli whole-body control, adds accelerated GR00T-RTC inference, and substantially improves training, simulation evaluation, environment setup, and reporting reliability. Highlights Add single- and dual-arm Franka support with ROS operators, a dedicated inference runner, PI0.5 QPos/ EEPose configurations, and an end-to-end deployment guide. Add Oli humanoid whole-body support with a whole-body operator, inference runner, GR00T finetuning configuration, tests, and usage documentation. Add accelerated GR00T-RTC inference with CUDA Graph-compatible prefix conditioning, previous-action prefill, partial action-chunk execution, and a UR3 deployment configuration. Expand LIBERO and RoboCasa evaluation with multi-task-per-GPU scheduling, resumable runs, rollout video export, result summarization, deterministic RoboCasa evaluation, and Feishu Sheets reporting. Simplify environment setup with one-command install and update scripts, split base/simulation/real-robot requirements, pinned simulation dependencies, and automated RoboCasa asset downloads. Modernize training configuration with registry-backed optimizers and LR schedulers, parameter-wise learning rates, stricter configuration validation, corrected step-based epoch accounting, and reliable DDP checkpoint restoration. Improve dataset and serving reliability with TorchCodec video decoding, opt-in dataset version validation, corrected distributed statistics aggregation, restored ARM/SARM LeRobot v3 loading, and full action-chunk preservation in ZMQ responses. Refresh Qwen3VL + GR00T LIBERO results, the RoboCasa deterministic 50-trial evaluation protocol, and documentation for Franka, Oli, UR3, data conversion, inference acceleration, and evaluation reporting. Upgrade Notes For an existing environment, run bash scripts/update_env.sh. For a fresh setup, use bash scripts/ install_env.sh sim-only, real-only, or full. Custom training configurations should migrate legacy runner.learning_rate, runner.weight_decay, runner.lr_scheduler_type, and runner.warmup_ratio fields to runner.optimizer and runner.lr_scheduler. requirements.txt now composes requirements-base.txt, requirements-sim.txt, and requirements- real.txt; rerun the setup or update script to install the pinned simulation stack. ZMQ clients now receive the complete denormalized action chunk without an extra batch dimension. Consumers that assumed a single action should update their shape handling.
If you use FluxVLA in your research, please cite it as below.
Vision-Language-Action Models (VLA), World-Action-Model (WAM), Robotics, Embodied AI
Vision-Language-Action Models (VLA), World-Action-Model (WAM), Robotics, Embodied AI
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
