arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

依意图行动:为视觉-语言-动作模型提炼行为意图

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Sangoh Lee, Sangwoo Mo, Wook-Shin Han

arXiv 2608.23478首次发表:更新:

发表机构

POSTECH; GSAI, POSTECH; IME, POSTECH(浦项科技大学; 浦项科技大学人工智能学院; 浦项科技大学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对视觉-语言-动作(VLA)模型动作解码器未明确行为语义目标的问题,提出INDI方法提炼行为级意图,在多个基准及真实任务中显著提升了GR00T-N1.7的性能。

AI 中文摘要

视觉-语言-动作(VLA)模型可将多模态上下文转化为机器人动作,但其动作解码器仍主要通过行为克隆进行训练。这一训练方式仅监督演示的运动指令,却未明确该行为在指令下的局部目标。基于未来的监督方法会用帧、潜在观测、轨迹或运动表征丰富动作学习,但这些信号捕捉的是可能发生事件的特定实现,而非即将执行行为的共享语义目标。本文提出意图提炼(INDI)方法,将行为级意图提炼至动作解码器中。训练期间,冻结的教师视觉-语言模型(VLM)会从当前观测、指令、粗略动作摘要及对应执行视频中解释演示片段。部署的VLA模型从其标准输入中恢复中间解码器层的多模态意图表征,并用该表征结合行为展开方式及达成目标的表征来组织动作预测。在SimplerEnv-Bridge上,INDI将GR00T-N1.7的性能从64.3%提升至84.7%;在RoboCasa Kitchen上,INDI将受控GR00T-N1.7基线的性能从64.1%提升至70.3%,且在两个基准上的π₀.₅指标均实现一致提升。在真实世界任务中,INDI将平均成功率从62.0%提升至68.7%,在长视程任务上的提升最高达12.0个百分点。进一步分析显示,恢复的潜在表征会被解码器使用,其捕捉了行为目标与执行进度,并以依赖目标的方式组织下游预测。这些结果表明,动作解码器可从显式建模其生成行为的语义目标中获益。

英文摘要

Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather than the shared semantic objective of the forthcoming behavior. We propose Intention Distillation (INDI), which distills behavior-level intent into the action decoder. During training, a frozen teacher VLM interprets a demonstrated segment from the current observation, instruction, coarse action summary, and corresponding execution video. From its standard inputs, the deployed VLA recovers the resulting multimodal intent representation at an intermediate decoder layer and uses it to organize action prediction together with representations of how the behavior unfolds and what it achieves. On SimplerEnv-Bridge, INDI improves GR00T-N1.7 from 64.3% to 84.7%, and on RoboCasa Kitchen it improves the controlled GR00T-N1.7 baseline from 64.1% to 70.3%, with consistent gains on $π_{0.5}$ across both benchmarks. In real-world tasks, INDI improves average success from 62.0% to 68.7%, with gains of up to 12.0 pp on longer-horizon tasks. Further analyses show that the recovered latent is used by the decoder, captures behavior objective and execution progress, and organizes downstream predictions in an objective-dependent manner. These results show that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.

CommentsProject page: https://leesangoh.github.io/indi-project-page/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑