arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01402cs.ROcs.AI

揭秘视觉-语言-动作模型在接触丰富任务中的失效时机、原因及修复方法

Demystifying When and Why VLAs Fail in Contact-Rich Tasks and How to Fix Them

Carlota Parés-Morlans, Nils Kuhn, Isabel Liu, Alberta Longhini, Jeannette Bohg

首次发表
浏览论文内容

中文总结 AI 辅助

该研究揭示视觉-语言-动作模型在接触丰富操纵任务中的两种失效模式,提出针对性机制整合而成的FACT模型,在5项任务上平均成功率达66%,优于现有基线。

中文摘要 AI 辅助

我们研究视觉-语言-动作(VLA)模型在需要精确物理交互的接触丰富操纵任务中表现不佳的时机与原因。现有研究主要通过力增强架构和训练时间正则化器解决接触失效,但失效的根本原因仍未被充分探索。我们识别出两种失效模式:精度失效源于流匹配策略训练不匹配,力失效源于力信号的独特结构。我们针对每种失效模式设计了针对性机制,将其整合为FACT模型。在涵盖近2500次真实世界rollout的评估中,FACT在5项接触丰富任务上的平均成功率达66%,而最佳现有基线仅为41%。

英文摘要

We address the problem of understanding when and why Vision-Language-Action models struggle with contact-rich manipulation tasks that require precise physical interaction. Prior work has primarily focused on addressing contact failures through force-augmented architectures and training-time regularizers, yet the root causes of these failures remain underexplored. We identify two distinct failure modes underlying this gap. Precision failures are rooted in a flow-matching policy training mismatch, and force failures arise from the distinctive structure of force signals. We address each failure mode with a targeted mechanism and combine them into FACT, which achieves 66% average success rate across five contact-rich tasks against 41% for the best prior baseline, in an evaluation spanning almost 2,500 real-world rollouts.

发表机构

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑