arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13225cs.CVcs.LGcs.RO

部件接地,而非动作知识:定位VLM可供性预测中的瓶颈

Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction

Sarthak Sattigeri

首次发表
浏览论文内容

中文总结 AI 辅助

该研究将可供性预测分解为部件识别与动作知识,通过实验证明部件接地是VLM的主要瓶颈,命名部件可显著提升动作准确率。

中文摘要 AI 辅助

基准测试一致认为,视觉语言模型对低级操作的推理能力较差,但总体准确率分数并不能说明是哪一步失败。我们将可供性问题混淆的两个步骤分开:识别应作用于物体的哪个部件,以及知道该部件需要什么动作。我们针对19个铰接物体,询问了来自三家开发者的八个模型,机器人应施加何种运动。在开放式提示下,当推(push)为正确答案时,64次评估中仅产生一次推,尽管推对于19个物体中的8个是正确的,并且每次都在提供的标签集中出现。检查输出揭示了原因:模型描述的是与被评分的部件不同的部件,例如解释如何拿起相机而不是按下其按钮。命名目标部件使每个模型的动作准确率提高了0.32至0.63,从0.158-0.474提高到0.684-0.947,推的召回率从0-1/8提高到7-8/8。在开放式提示下,没有模型能胜过忽略图像的恒定答案;一旦部件被命名,所有八个模型都能胜过。当要求用自由文本描述同一部件且无标签集时,模型对8个中的6至8个产生了按压语言。这些结果很难与缺失动作知识相协调,反而指向部件接地作为主要瓶颈,这一模式在所有三个模型家族中均成立,且不随能力增强而减弱。命名部件提供了接地变量,因此这界定了完美部件检测器所能提供的上限,而非展示通用的力学模型。两个支持性结果一致:在真实照片上,八个模型中只有三个的抓取点定位优于恒定基线,而在渲染物体上则没有模型优于基线。我们还记录了我们自己的两个测量错误:一个阈值使恒定基线得分为0.929,以及一个标签规则在19个物体中的4个上出错,这两个错误都是通过将我们的数字与平凡替代方案进行测试才发现的。

英文摘要

Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fails. We separate two steps that affordance questions conflate: identifying which part of an object to act on, and knowing what action that part requires. Across 19 articulated objects we asked eight models, spanning three developers, what motion a robot should apply. Under an open prompt, push was produced once in 64 evaluations where it was correct, despite being correct for 8 of 19 objects and appearing in the offered label set every time. Inspecting the outputs showed why: models described a different part than the one being scored, e.g. explaining how to pick up a camera rather than press its button. Naming the target part raises action accuracy by 0.32 to 0.63 for every model, from 0.158-0.474 to 0.684-0.947, and push recall from 0-1/8 to 7-8/8. No model beats a constant answer that ignores the image under the open prompt; once the part is named, all eight do. Asked to describe the same part in free prose with no label set, models produce pressing language for 6 to 8 of 8. These results are hard to reconcile with missing action knowledge, and instead point to part grounding as the dominant bottleneck, a pattern that holds across all three model families and does not diminish with capability. Naming the part supplies the grounding variable, so this bounds what a perfect part detector would offer rather than demonstrating a general model of mechanics. Two supporting results agree: on real photographs only three of eight models localize grasp points better than a constant baseline, and on rendered objects none do. We also document two measurement errors of our own, a threshold that let a constant baseline score 0.929 and a labelling rule wrong on 4 of 19 objects, both caught only by testing our numbers against trivial alternatives.

发表机构

  • Manipal University Jaipur(马尼帕尔大学斋浦尔分校)

机构由 AI 辅助整理,请以论文原文为准。

↑