VANE:通过未来视觉表示预测实现视觉-语言-动作模型的可靠测试时训练
VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction
浏览论文内容
中文总结 AI 辅助
本文提出VANE框架,通过条件提示自适应、未来视觉后果学习及候选更新隔离评估的方法,提升VLA策略在闭环操纵中的可靠性,在SimplerEnv WidowX上将平均成功率提高3.2个百分点。
中文摘要 AI 辅助
测试时训练(TTT)提供了一种轻量级方法,可从未标记的部署流中自适应视觉-语言-动作(VLA)策略,但在闭环操纵中可靠使用仍存在困难:共享自适应空间可能混合不兼容的任务修正,而在线更新可能在其后果已知前改变后续动作。本文提出用于VLA策略的可靠TTT框架VANE,该框架将提示自适应条件设置为当前视觉-语言上下文,并从已执行动作的未来视觉后果中学习。候选更新与实时策略隔离,通过后续观测评估,仅在有未来证据支持时才提交,使自适应具有选择性和可逆性。在SimplerEnv WidowX上,VANE较对应TTT基线将平均成功率提高了3.2个百分点;在Google Robot上的结果进一步表明,部署时的增益仍依赖于任务和 embodiment。这些结果共同证明了一种在交互过程中自适应VLA策略的受约束、基于证据的方法。
英文摘要
Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by $3.2$ percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Li Auto Inc.(理想汽车)
机构由 AI 辅助整理,请以论文原文为准。