arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34175cs.RO

FailPatch:视觉-语言-动作模型的失败残差修补

FailPatch: Failure Residual Patching for Vision-Language-Action Models

  • Dexmal
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

Peng Yu, Jiacheng Wang, Ziheng Zhang, Xuchong Zhang, Baoting Li, Zhuoyuan Yu, Yuxiang Chen, Tiancai Wang, Hongbin Sun

AI总结:

FailPatch提出失败驱动的残差修补框架,通过零门控残差专家库在冻结VLA策略中解耦动作与可靠性监督,仅用0.52%参数提升成功率11.0至16.7个百分点。

AI中文摘要:

视觉-语言-动作(VLA)策略通常使用成功演示进行适配,这些演示提供了直接的动作监督,但很少覆盖易失败的状态。部署过程中的失败暴露了这些状态,却缺乏传统监督学习所需的纠正动作。我们提出了FailPatch,一个失败驱动的残差修补框架,将动作监督与执行可靠性监督解耦。成功演示为策略应如何行动提供基础,而部署轨迹则指示其行为何时变得不可靠。我们进一步观察到,动作隐藏表示在可靠状态与失败关联状态之间表现出清晰的线性可分性,同时直接调节动作生成。基于这些见解,FailPatch在冻结的VLA策略的动作隐藏空间中引入了一个零门控残差专家库。统一的保留-重定向-信任目标在可靠状态下保留原始策略,在失败关联状态下选择残差专家,并在有界干预下将表示从失败区域重定向到成功关联区域。仅用0.52%的可训练参数,FailPatch在四个长时程RoboTwin任务的干净评估中将成功率提高了11.0个百分点,在干净到随机泛化中提高了9.5个百分点,并在三个真实世界任务中比基线提高了16.7个百分点。项目和代码:此https URL。

英文摘要:

Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose FailPatch, a failure-driven residual patching framework that decouples action supervision from execution-reliability supervision. Successful demonstrations ground how the policy should act, while deployment trajectories indicate when its behavior becomes unreliable. We further observe that action hidden representations exhibit clear linear separability between reliable and failure-associated states while directly conditioning action generation. Building on these insights, FailPatch introduces a Null-gated Residual Expert Bank into the action hidden space of a frozen VLA policy. A unified Preserve--Redirect--Trust objective retains the original policy in reliable states, selects residual experts in failure-associated states and redirects representations from failure regions toward success-associated regions under bounded intervention. With only 0.52% trainable parameters, FailPatch improves success rates by 11.0 percentage points on four long-horizon RoboTwin tasks under clean evaluation, 9.5 percentage points under clean-to-random generalization, and 16.7 percentage points over the baseline across three real-world tasks. Project and code: https://github.com/yupeng-2003/FailPatch.

↑