预测监督的方式决定VLA策略学习的内容
Where Predictive Supervision Goes Shapes What VLA Policies Learn
浏览论文内容
中文总结 AI 辅助
本文研究预测监督如何影响VLA策略的视觉表征,发现不同预测接口导致表征差异,直接场景匹配的未来监督可提升策略在分布偏移下的鲁棒性。
中文摘要 AI 辅助
未来预测越来越多地被用于改进视觉-语言-动作(VLA)策略,其前提是预期场景演变有助于产生对控制有用的表征。然而,仅凭预测质量并不能证明策略已学习到更好的动作表征。在分布偏移下,这种区别至关重要,因为成功的控制依赖于在熟悉配置之外保持空间状态和可能的场景变化。我们研究了决定预测监督是否改善VLA策略所用视觉表征的因素。通过使用匹配的目标构建、预测视界和训练条件的受控比较,我们发现不同的预测接口会产生显著不同的预测和视觉表征,包括在空间、动态和动作信息方面,这些信息会迁移到熟悉场景之外。我们将这些差异追溯到预测误差如何塑造策略的视觉流。与这一受控发现一致,使用更直接、场景匹配的未来监督训练的VLA策略在模拟和物理分布偏移下表现出更强的鲁棒性。总之,我们的结果将未来预测框定为表征学习的设计问题,其对控制的价值取决于监督是否到达策略行动所依赖的表征。
英文摘要
Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy's visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.
发表机构
- Graduate School of Data Science(数据科学研究生院)
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。