arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习在视觉中断下使用视觉-语言-动作模型行动

Learning to Act under Visual Interruptions with Vision-Language-Action Models

Mingle Jiang, Rui Xu, Yunke Wang, Chang Xu

arXiv 2609.35003首次发表:更新:

发表机构

The University of Sydney(悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对相机中断导致视觉缺失的问题,提出MAIL-Bench基准和MINT方法,通过训练与选择性补充观测,显著提升VLA模型在相机丢失下的操作成功率。

AI 中文摘要

视觉-语言-动作(VLA)模型在机器人操作中展现了强大的能力,但它们通常在任务执行过程中所有相机流均可用的情况下进行开发和评估。当相机在任务执行期间停止传送帧时,策略必须在缺少来自缺失视图的后续观测的情况下继续行动。尽管这一问题具有实际重要性,但此类中断如何影响闭环操作仍未得到充分理解。为了研究这一问题,我们引入了MAIL-Bench,一个使用VLA模型评估视觉中断的基准。通过在每个策略成功的参考轨迹的多个阶段中断不同相机,MAIL-Bench衡量了当视觉输入变得不可用时策略保持其能力的程度。基于这一基准,我们提出了MINT,它首先训练VLA策略在缺失视觉输入的情况下保持功能。在推理时,MINT使用光流外推或动作条件世界模型选择性地补充缺失的观测,并在预测视图变得不可靠时撤回它们。在π0.5和GR00T N1.5上的实验表明,MINT在相机丢失下显著提高了任务成功率,优于原始模型。在AgiBot G2上的实验进一步展示了在相机丢失下的真实机器人部署。该基准可在以下网址获取:https://this https URL

英文摘要

Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting without access to subsequent observations from the missing view. Despite its practical importance, how such interruptions affect closed-loop manipulation remains insufficiently understood. To investigate this problem, we introduce MAIL-Bench, a benchmark that evaluates visual interruptions with VLA models. By interrupting different cameras at multiple stages of each policy's successful reference trajectory, MAIL-Bench measures how well policies retain their capabilities when visual inputs become unavailable. Building on this benchmark, we propose MINT, which first trains VLA policies to remain functional under missing visual inputs. At inference time, MINT selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable. Experiments on $π_{0.5}$ and GR00T N1.5 show that MINT significantly improves task success under camera loss over the original models. Experiments on AgiBot G2 further demonstrate the real-robot deployment under camera loss. The benchmark is available at https://minglejiang.github.io/Mail-Bench/

Commentshttps://minglejiang.github.io/Mail-Bench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑