Action- and World Model-based Rewards Improve Vision-Language-Action Model Post-Training
| Title | Action- and World Model-based Rewards Improve Vision-Language-Action Model Post-Training |
| Publication Type | Conference Paper |
| Year of Publication | 2026 |
| Authors | Hung C-Y, Majumder N, Deng H, Renhang L, Ang Y, Herremans D., Wang Z, Poria S. |
| Conference Name | NeurIPS workshop: Robot Learning with World Models: Capabilities, Frontiers, and Challenges |
| Abstract | Vision-language-action (VLA) models have shown strong performance across embodied AI tasks, but they remain insufficiently reliable for real-world deployment. Post-training with reward signals has proven effective for improving generative models across language, vision, and audio. Motivated by this, we propose various reward design for post training VLA. We design two complementary rewards for constructing preference data: a world-model-based reward that estimates whether candidate actions move toward a specified goal or subgoal, and an action-distance reward that regularizes actions toward feasible expert trajectories. Using preference pairs generated from these rewards, we apply Direct Preference Optimization (DPO) to adapt NORA-1.5. Experiments on both simulated and real-world benchmarks show consistent improvements over supervised fine-tuning baselines. These results suggest that proxy reward design offers a scalable path for improving VLA models without costly online robot rollouts. |
| URL | https://openreview.net/forum?id=DSGKJNgIzn |