arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03611cs.ROcs.AI

FailBench:视觉语言模型(VLMs)在判断机器人任务成功方面的可靠性如何?

FailBench: How Reliable are VLMs at Judging Robot Task Success?

Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出机器人故障检测基准FailBench,评估13种VLM检测器,发现微调故障检测的VLM表现弱于通用VLM,输入级干预可提升最佳检测器性能2.4个百分点。

中文摘要 AI 辅助

视觉语言模型(VLMs)正越来越多地被用于评估机器人操作结果,但现有基准提供的跨域泛化证据有限。我们推出FailBench,一个用于机器人故障检测的基准,包含来自14个公开来源(12个真实世界、2个模拟)的2197次操作尝试。在FailBench中,75%的故障是自然发生的,且6个真实世界来源来自非故障检测数据集。对13种基于VLM的检测器进行评估后,我们发现最佳模型仅达到0.77的平均平衡准确率。值得注意的是,为故障检测微调的模型始终表现不如通用VLMs及其自身的预训练基线。性能高度依赖于所需的视觉证据:当结果取决于可观察的物体运动时,模型接近饱和,但在接触密集的装配任务上,性能降至接近随机水平(平衡准确率<0.60)。误差分析显示存在系统性偏差,即在模糊证据下倾向于预测成功,即使增加推理努力,这种偏差仍然存在。最后,我们表明输入级干预——对结果相关区域进行空间定位和裁剪——可使最佳检测器提升2.4个百分点,且无需额外训练。

英文摘要

Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benchmarks offer limited evidence of cross-domain generalization. We introduce FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated). In FailBench, 75% of failures occur naturally, and six real-world sources come from non-failure-detection datasets. Evaluating 13 VLM-based detectors, we find the best model achieves only 0.77 mean balanced accuracy. Notably, models fine-tuned for failure detection consistently underperform general-purpose VLMs and their own pretrained baselines. Performance depends heavily on required visual evidence: models approach saturation when outcomes depend on observable object motion, but degrade to near-chance (<0.60 balanced accuracy) on contact-intensive assembly tasks. Error analysis reveals a systematic bias toward predicting success under ambiguous evidence, which persists even with increased reasoning effort. Finally, we show that input-level intervention--spatially localizing and cropping outcome-relevant regions--improves the top detector by 2.4 percentage points without extra training.

↑