发表机构
ISIR, Sorbonne Université(法国国家信息与自动化研究所,索邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对基于流匹配的视觉-语言-动作模型细粒度行为控制不足的问题,提出DiMaS分布匹配引导策略,能有效控制行为,研究其泛化性并分析原因,推动了视觉运动设置下的分布匹配设计。
AI 中文摘要
基于流匹配的视觉-语言-动作(VLA)模型已成为机器人操作的强大策略,但细粒度行为控制这一关键能力仍未得到充分探索。表示引导是语言和视觉-语言模型的既定可解释性工具,不过在VLA中这些经典方法存在不足。我们提出了DiMaS,一种针对流匹配VLA的分布匹配引导策略,它在表示分布之间传输而非沿固定方向移动,并证明其能有效控制两个先进VLA的行为。我们进一步研究该策略在学习和评估任务差异增大时的泛化性,分析动作专家表示结构,解释经典线性引导在视觉运动设置中不足的原因,即行为特征可线性解码但不可线性引导,这推动了DiMaS的分布匹配设计。代码可公开获取。
英文摘要
Flow-matching-based vision-language-action (VLA) models have emerged as powerful policies for robotic manipulation, yet a critical capability remains underexplored: fine-grained behavioral control, the ability to govern how a robot performs a task by intervening on its internal representations. Representation steering is a well-established interpretability tool for language and vision-language models, where behavioral features are typically encoded as linear directions, but we show that these classic methods fall short in VLAs. We propose DiMaS, a Distribution-Matching Steering strategy tailored to flow-matching VLAs, which transports between representation distributions rather than shifting along a fixed direction, and show that it effectively controls behavior across two state-of-the-art VLAs. We further examine the generalizability of this strategy as the tasks it is learned from and evaluated on grow increasingly dissimilar, characterizing where behavioral control transfers and where it weakens. Finally, through an analysis of the representation structure of the action expert, we explain why classical linear steering falls short in the visuomotor setting: behavioral features are linearly decodable but not linearly steerable, which motivates the distribution-matching design of DiMaS. Our code is publicly available at https://github.com/pegah-kh/dimas, with additional results and videos at https://pegah-kh.github.io/dimas/