arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

纠正WHERE,保留HOW:通过参照引导实现视觉-语言-动作模型的组合泛化

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin

arXiv 2609.38616首次发表:更新:

发表机构

Case Western Reserve University(凯斯西储大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ReGuide,一种无需训练的包装器,通过语义和几何重新绑定引导末端执行器,提升VLA模型在组合变化下的成功率,同时保持标准任务性能。

AI 中文摘要

尽管视觉-语言-动作(VLA)模型能够实现灵活的动作生成,但其在多样化环境要素(包括操作对象、目的地和背景)上的泛化能力受到机器人训练数据缺乏多样性的限制。在此类数据上端到端训练的VLA模型倾向于利用视觉捷径,将动作与任务无关的视觉特征关联起来,而非预期的任务语义。这些捷径阻碍了策略对已见元素的重新组合,即组合泛化。现有方法通过任务相关感知或针对性的数据多样化来缓解这种纠缠,但未提供针对未见重组的显式机制,并且需要特定骨干网络的修改和重新训练。我们观察到,在这种重组下,VLA模型往往在全局定位上失败,但在熟悉配置中保留了接近正确目标的局部操作技能。因此,我们提出参照引导(ReGuide),一种无需训练的包装器,它利用来自定位模块的对象姿态,结合语义和几何重新绑定,将末端执行器引导至被指示参照物的演示支持配置中,使冻结策略能够恢复执行。在多个VLA骨干网络的模拟实验以及真实机器人上的实验表明,ReGuide在组合变化下的成功率分别提高了最多56.8和75.0个百分点,同时保持了标准任务的性能。

英文摘要

While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including manipulated objects, destinations, and backgrounds, is limited by the lack of diversity in robotic training data. Trained end-to-end on such data, VLAs tend to exploit visual shortcuts, associating actions with task-irrelevant visual features rather than the intended task semantics. These shortcuts block recomposition of elements already seen by the policy, that is, compositional generalization. Existing approaches mitigate such entanglement through task-relevant perception or targeted data diversification, but offer no explicit mechanism for unseen recomposition and require backbone-specific modifications with retraining. We observe that under such recomposition, VLAs often fail at global grounding while retaining local manipulation skills that recover near the correct target in familiar configurations. Therefore, we propose Referential Guidance (ReGuide), a training-free wrapper that, given object poses from a grounding module, combines semantic and geometric rebinding to guide the end-effector into demonstration-supported configurations of the instructed referent, where the frozen policy can resume execution. Experiments in simulation across multiple VLA backbones as well as on a real robot show that ReGuide improves success rates under compositional shifts by up to 56.8 and 75.0 percentage points, respectively, while preserving standard-task performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑