arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31792cs.AIcs.RO

ConflictVLA-Bench:视觉-语言-行动模型对前提冲突的行为响应基准测试

ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts

Liyu Hou, Yuan Wu, Yi Chang

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA模型对无效前提的响应,提出ConflictVLA-Bench基准,通过冲突与参考轨迹配对评估,发现模型普遍存在“失败持续性”,表明仅凭结果评估不足。

中文摘要 AI 辅助

尽管视觉-语言-行动(VLA)模型在操作任务上表现强劲,但它们对无效任务前提的响应仍未得到充分探索。现有对前提冲突的评估往往关注最终任务结果,然而任务失败本身无法区分行为脱离与持续追求后执行错误,我们将后一种模式称为“失败持续性”。为研究这一现象,我们引入了ConflictVLA-Bench,它将冲突轨迹与前提一致的参考轨迹配对,并同时评估结果和执行过程。该基准基于LIBERO构建,包含2,826个提示条件下的冲突任务,涵盖四个冲突家族、四种结构配置和两种提示条件。在所有八个VLA模型中,无效前提使原始目标完成率至少下降17.3个百分点,其中OpenVLA的下降幅度达到56.2个百分点。关键的是,即使模型在前置一致任务上成功而在匹配的冲突任务上失败,它们也常常继续接近原始目标,保留早期轨迹结构,并表现出有限的动作幅度抑制。因此,“失败持续性”在评估的模型中反复出现。显式前提检查并不能一致地产生选择性和协调的行为变化。这些发现表明,最终失败本身既不能确定行为脱离,也不能确定拒绝,仅凭结果不足以进行VLA评估。实验数据和更多细节可在项目页面获取:此 https URL

英文摘要

While Vision-Language-Action (VLA) models perform strongly on manipulation tasks, their responses to invalid task premises remain underexplored. Existing evaluations of premise conflicts often focus on terminal task outcomes, yet task failure alone cannot distinguish behavioral disengagement from continued pursuit followed by an execution error. We call the latter pattern Failed Persistence. To study this phenomenon, we introduce ConflictVLA-Bench, which pairs conflict rollouts with premise-consistent reference rollouts and evaluates both outcomes and execution processes. Built on LIBERO, the benchmark contains 2,826 prompt-conditioned conflict tasks spanning four conflict families, four structural configurations, and two prompt conditions. Across all eight VLAs, invalid premises reduce original goal completion by at least 17.3 percentage points, with the reduction reaching 56.2 percentage points for OpenVLA. Crucially, even when models succeed on premise-consistent tasks and fail on their matched conflict tasks, they often continue to approach the original targets, retain early trajectory structure, and show limited action magnitude suppression. Failed Persistence therefore recurs across the evaluated models. Explicit premise checking does not consistently produce selective and coordinated behavioral changes. These findings show that terminal failure alone establishes neither behavioral disengagement nor refusal and that outcomes alone are insufficient for VLA evaluation. Experimental data and additional details are available on the project page: https://github.com/EmbodiedAISurvey/ConflictVLA-Bench

补充信息

↑