arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06114cs.RO

成功何处失效:面向鲁棒视觉-语言-动作模型的失败边界学习

Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models

Yanzhe Chen, Zhijun Cao, Mike Zheng Shou

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型微调忽略失败边界的问题,提出DLS方法,通过数字孪生回放发现边界、语义进度定位和方向性塑造,提升真实机器人操作的鲁棒性。

中文摘要 AI 辅助

通过监督微调(SFT)适配的视觉-语言-动作(VLA)模型继承了一种结构性不对称:专家示范教会了策略成功行为所在之处,但未提供关于其在何处不再可靠的信号。我们认为,鲁棒的VLA适配因此不应被视为进一步的示范拟合,而应被视为**失败边界学习**——即*发现*、*定位*和*塑造*可恢复偏差与任务失败之间边界的问题。为实例化这一观点,我们提出**DLS**:基于来自少量真实示范和模拟协同训练构建的**真实接地行为先验**,DLS通过在线策略数字孪生回放大规模*发现*失败边界。并非将每次回放简化为二元标签,**语义进度定位**利用特权模拟器状态分配进度感知信号,以捕获失败边界在*何处*被跨越,而不仅仅是*是否*被跨越。这些信号驱动流动力学中的**方向性边界塑造**——强化产生成功的去噪方向并抑制产生失败的方向,无需动作似然或辅助评论器。在真实机器人操作任务中,DLS相比SFT和在线强化学习基线提升了鲁棒性,尤其在随机初始状态和未见视觉条件下。

英文摘要

Vision-language-action (VLA) models adapted through supervised fine-tuning (SFT) inherit a structural asymmetry: expert demonstrations teach the policy where success behavior lies, but provide no signal about where it ceases to be reliable. We argue that robust VLA adaptation should therefore be viewed not as further demonstration fitting, but as Failure-Boundary Learning -- the problem of Discovering, Localizing, and Shaping the boundary between recoverable deviations and task failure. To instantiate this view, we propose DLS: built on a real-grounded behavioral prior from few real demonstrations and simulated co-training, DLS discovers failure boundaries at scale through on-policy digital twin rollouts. Rather than reducing each rollout to a binary label, semantic progress localization uses privileged simulator states to assign progress-aware signals that capture where the failure boundary is crossed, not merely whether. These signals drive directional boundary shaping in the flow dynamics -- reinforcing success-producing denoising directions and suppressing failure-producing ones, without action likelihoods or auxiliary critics. Across real-robot manipulation tasks, DLS improves robustness over SFT and online RL baselines, especially under randomized initial states and unseen visual conditions.

发表机构

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑