发表机构
The Chinese University of Hong Kong; Centre for Perceptual and Interactive Intelligence; Duke University; Sun Yat-sen University; TengenX; Tencent Robotics X; Shenzhen Loop Area Institute; Guangdong Key Laboratory of Big Data Analysis and Processing(香港中文大学; 感知与交互智能中心; 杜克大学; 中山大学; 天工科技; 腾讯Robotics X实验室; 深圳河套学院; 广东省大数据分析与处理重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言机器人操纵中VLAs无法从执行错误恢复的问题,提出含“重试”“重置”范式的FLARE框架,结合MLLM实现故障恢复,显著提升接触密集型操纵任务的成功率与鲁棒性。
AI 中文摘要
视觉-语言-动作模型(VLAs)在泛化到复杂、长时程机器人操纵任务方面展现出巨大潜力,但其性能仍较为脆弱,因为它们通常在轨迹单调、无故障的演示数据上训练。这种对“完美”数据的依赖使得VLAs无法从常见的执行错误中恢复,比如抓取失败、物体掉落或意外碰撞。本文提出一种新颖的框架FLARE,通过“重试(Retry)”与“重置(Reset)”范式赋予VLAs强大的错误恢复能力。首先,我们引入“重试”机制,通过注入扰动和将机器人位姿与环境状态解耦的桥接段融入演示数据,使策略能自主处理执行偏差。其次,为解决关键的状态破坏型(OOD)故障,我们设计了“重置”流水线:利用多模态大语言模型(MLLM)进行离线故障分析,从执行视频中自动识别OOD状态;该分析支持高效、有针对性地收集以物体为中心的小型“重置”技能库,这些技能经训练可将环境恢复至任务有效状态。我们的完整框架整合了这些学习到的策略,推理阶段由在线MLLM监视器在任务执行与“重置”技能间进行仲裁。在具有挑战性的接触密集型操纵任务上的实验表明,我们的方法显著提升了任务成功率与鲁棒性。
英文摘要
Vision-Language-Action Models~(VLAs) have demonstrated significant promise in generalizing to complex, long-horizon robotic manipulation tasks. However, their performance remains brittle, as they are typically trained on trajectory-monotonic, failure-free demonstrations. This reliance on ``perfect" data leaves them unable to recover from common execution errors, such as a missed grasp, a dropped object, or an unexpected collision. In this paper, we propose FLARE, a novel framework that endows VLAs with robust error recovery capabilities through a ``Retry" and ``Reset" paradigm. First, we introduce a ``Retry" mechanism by injecting perturbation and bridging segments that decouple robot pose from environment state into demonstrations, enabling the policy to autonomously handle execution deviations. Second, to address critical, state-breaking (OOD) failures, we introduce a ``Reset" pipeline. We leverage an MLLM for offline failure analysis to automatically identify OOD states from execution videos. This analysis enables the efficient, targeted collection of a small library of object-centric ``Reset" skills, which are trained to restore the environment to a task-valid state. Our full framework integrates these learned policies. At inference, an online MLLM monitor arbitrates between task execution and ``Reset" skills. Experiments on challenging, contact-rich manipulation tasks show our approach significantly improves task success and robustness.
CommentsAccepted to CVPR 2026