arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39820cs.ROcs.AI

从运行时反馈中通过失败库自进化学习:面向视觉-语言-动作模型

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型在复杂环境中策略与防护不匹配的问题,提出FailBank四阶段自进化框架,将运行时反馈转化为持续策略改进,在VLA-Arena上显著提升成功率并降低成本。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型在机器人操作任务中具有广泛的泛化能力,但复杂环境需要在任务成功与意外接触之间取得平衡。运行时防护可以纠正单个动作,但底层策略保持不变,因此重复的分歧可能造成持续的策略-防护不匹配,阻碍任务进展。为解决这一挑战,我们提出了FailBank,一个四阶段自进化框架,将运行时反馈转化为持续的策略改进。在收集阶段,一个固定的基于CBF的安全模块作为仅观察的教师,产生反事实修正,而策略保持控制。结果感知接纳随后将有用的提议转化为修正目标,并保留成功的未修正动作作为静默锚点,用于受保护的LoRA更新。我们在VLA-Arena基准上,跨两个难度级别和两个VLA骨干网络评估了FailBank。与基础策略相比,FailBank改善了联合成功-成本操作点。在两个骨干网络上,FailBank分别将任务成功率提高了8.5和6.9个百分点,同时将策略引发的累积成本分别降低了35.6%和23.8%。与运行时防护相比,FailBank将任务成功率提高了25.4和9.5个百分点,同时保持可比的策略引发累积成本。这些结果表明,运行时反馈可以作为持续的策略监督,而不仅仅是临时的动作约束。

英文摘要

Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress. To address this challenge, we introduce FailBank, a four-stage self-evolving framework that converts runtime feedback into persistent policy improvement. During collection, a fixed CBF-based safety module serves as an observe-only teacher, producing counterfactual corrections while the policy remains in control. Outcome-aware admission then converts useful proposals into corrective targets and retains successful uncorrected actions as quiet anchors for guarded LoRA updates. We evaluate FailBank on the VLA-Arena benchmark across two difficulty levels and two VLA backbones. Compared with the base policies, FailBank improves the joint success-cost operating point. Across the two backbones, FailBank improves task success rate by 8.5 and 6.9 percentage points, while reducing policy-induced cumulative cost by 35.6\% and 23.8\%, respectively. Compared with runtime shielding, FailBank raises task success rate by 25.4 and 9.5 percentage points, while maintaining comparable policy-induced cumulative cost. These results show that runtime feedback can serve as persistent policy supervision rather than only as a temporary action constraint.

发表机构

  • University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑