arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分离视觉-语言模型中的感知与推理:晶体结构的无模型渲染上限

Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures

Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban

arXiv 2609.00663首次发表:更新:

AI 中文总结

该研究提出无模型渲染上限方法,可分离视觉-语言模型的感知与推理错误,在2160个晶体结构上验证其有效性,发现无语言监督视觉模型性能优于14个视觉-语言模型,为基准构建提供规则。

AI 中文摘要

多模态评估无法判断视觉-语言模型是误读了图像还是对图像的推理出错,因为现有的所有用于区分这两种情况的方法都会在流程中引入第二个模型。我们提出了渲染上限(render ceiling),这是一种通过渲染已知对象构建的无模型基准参考:反转冻结的相机并重新求解跨视角对应关系,可精确恢复图像所支持的答案。我们证明该上限仅会因一组可枚举的投影巧合而失效,并在2160个渲染的晶体结构上验证该集合为空,因此模型的每一处性能缺陷都属于模型本身。在14个视觉-语言模型上的实验显示,将精确几何信息作为文本提供给所有模型后,每个模型的性能都有所提升,但其中13个模型仅弥补了不到一半的性能差距;而一个无语言组件的监督视觉模型在读取相同图像时准确率达到0.8952,超过了所有视觉-语言模型。该工具可揭示下游准确率会误归因于推理的提取阶段造假,为基准构建者提供相机放置规则,并可迁移到任何具有可逆前向渲染的基准。

英文摘要

Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separating the two places a second model in the loop. We introduce the render ceiling, a model-free reference for benchmarks built by rendering known objects: inverting the frozen cameras and re-solving cross-view correspondence recovers exactly the answer the images support. We prove the ceiling fails only through an enumerable set of projection coincidences and certify that set empty on 2,160 rendered crystal structures, so every point of a model's deficit belongs to the model. Across fourteen vision-language models, supplying exact geometry as text lifts every model yet closes under half the gap for thirteen, while a supervised vision model with no language component reads the same images at 0.8952, above every vision-language model. The instrument exposes extraction-stage fabrication that downstream accuracy would misattribute to reasoning, yields camera-placement rules for benchmark builders, and transfers to any benchmark with an invertible forward rendering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑