arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LIBERO-VIFO:评估视觉-语言-动作模型视觉线索跟随的能力与安全性的基准

LIBERO-VIFO: Benchmarking the Capability and Safety of Visual Cue Following in Vision-Language-Action Models

Zhengyan Qian, Rui Yan, Alex Jinpeng Wang, Jinhui Tang

arXiv 2608.17600首次发表:更新:

AI 中文总结

研究人员提出LIBERO-VIFO基准,评估7种VLA模型的视觉线索跟随能力与安全性,发现当前VLA存在无语言指令时执行线索指示任务的未授权跟随风险,确立了以视觉为中心的安全新视角。

AI 中文摘要

视觉线索正越来越多地被用于指导机器人学习,但视觉-语言-动作(VLA)模型能否可靠跟随授权线索同时忽略未授权线索仍不明确。现有研究仅覆盖范围狭窄的线索形式,且聚焦于最终任务成功,仅能对线索跟随能力进行粗略评估;将所有视觉线索视为授权线索的做法也未探索未授权跟随的安全风险。为解决这些差距,我们提出LIBERO-VIFO,一个用于评估VLA模型视觉线索跟随能力与安全性的基准。LIBERO-VIFO定义了涵盖不同形式的8类视觉线索,共设置两部分的4种协议:第一部分测试线索理解与授权跟随,第二部分评估语言-线索冲突及空语言条件下的未授权视觉线索跟随。对7种VLA模型的评估显示,尽管视觉线索理解无法可靠转化为执行结果,但当前VLA模型能够在无语言指令的情况下执行线索指示的任务,暴露出未授权视觉线索跟随的新兴风险。对场景实例化线索、安全关键设置及真实机器人部署的扩展实验证实了这些发现。LIBERO-VIFO将视觉线索跟随的能力与安全性纳入系统评估,为VLA领域确立了以视觉为中心的安全这一新视角。

英文摘要

Visual cues are increasingly adopted to guide robot learning, but whether Vision-Language-Action (VLA) models can reliably follow authorized cues while disregarding unauthorized ones remains unclear. Existing work covers only a narrow range of cue forms and focuses on final task success, providing only a coarse assessment of cue-following capability. Treating all visual cues as authorized also leaves safety risks of unauthorized following unexplored. To address these gaps, we introduce LIBERO-VIFO, a benchmark to evaluate both the capability and safety of visual cue following in VLA models. LIBERO-VIFO defines eight visual cue families spanning diverse forms. A total of four protocols in two parts are defined: Part I tests cue understanding and authorized following, while Part II evaluates unauthorized visual cue following under language-cue conflict and empty language conditions. Evaluating seven VLA models reveals that although visual cue understanding does not reliably translate into execution, current VLAs are able to execute cue-indicated tasks without language instruction, exposing an emerging risk of unauthorized visual cue following. Extended experiments on scene-instantiated cues, safety-critical settings, and real-robot deployment corroborate these findings. LIBERO-VIFO brings both the capability and safety of visual cue following into systematic evaluation, establishing visual-centric safety as a new perspective for the VLA community.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑