AI 中文总结
本研究针对VLA模型训练中视觉语言高级表示与机器人低级动作的不匹配问题,提出UVT统一视觉运动目标,无需改动架构或额外数据,在仿真与真实双臂任务中提升了VLA的训练效率、性能与鲁棒性
AI 中文摘要
视觉语言智能体(VLA)模型经过训练,可从视觉和语言观测中预测机器人动作,这是一种自然选择,但存在不匹配:视觉语言模型(VLM)编码场景和目标的丰富高级表示,而机器人动作是任务结构有限的低级信号。我们探究改变策略的训练预测内容而非其架构设计,能否生成更优且训练更高效的策略。我们提出UVT(Unified Visuomotor Target,统一视觉运动目标),一种联合编码运动控制和视觉场景转换信息的统一潜在预测目标,无需架构改动和额外数据。将其应用于两个代表性VLA系统,跨越仿真基准和真实双臂操作任务,UVT提升了训练效率、最终任务性能及策略鲁棒性,在有限训练预算和具挑战性环境条件下增益尤为显著。演示视频和更多定性结果可在项目网页获取:this https URL
英文摘要
VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representations of scenes and goals, while robot actions are low-level signals with limited task structure. We ask whether changing what the policy is trained to predict, rather than how it is architecturally designed, can yield better and more efficiently trained policies. We propose UVT (Unified Visuomotor Target), a unified latent prediction target that jointly encodes motor control and visual scene transition information, requiring no architectural changes and no additional data. Applied to two representative VLA systems across simulation benchmarks and real bimanual manipulation tasks, UVT improves training efficiency, final task performance, and policy robustness, with particularly strong gains under limited training budgets and challenging environmental conditions. Rollout videos and additional qualitative results are available at our project webpage: https://unified-visuomotor-targets.github.io/
CommentsAccepted at IROS 2026. Project page: https://unified-visuomotor-targets.github.io/