arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉模型学习的是物理约束还是渲染捷径?一个用于基础物理一致性的反事实基准

Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency

M. Moein Esfahani, Sepehr Salem, Mohammed Alser, Vince Calhoun

arXiv 2610.09205首次发表:更新:

发表机构

Tri-institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS); Georgia State University; Georgia Institute of Technology; Emory University(转化神经影像与数据科学三机构中心; 佐治亚州立大学; 佐治亚理工学院; 埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出反事实基准,用合成图像检测编辑场景的物理违规,发现视觉模型高准确率部分源于渲染捷径而非物理理解。

AI 中文摘要

现代图像编辑模型可以在满足文本指令的同时破坏编辑场景的物理规律。新物体可能不投射阴影,镜子可能无法反射可见几何体,或者物体可能漂浮在应支撑它的表面之上。我们研究物理合理性诊断,即检测编辑后的图像是否违反场景物理,命名违规类型,定位受影响区域,并用语言解释失败原因。我们引入一个反事实基准,其受控合成部分使用Mitsuba 3从500个场景族生成5500张图像。每个场景族包含一张干净图像和十个匹配的违规图像,涉及阴影、反射、支撑、表面响应和遮挡。渲染器流程提供类别标签、受影响区域掩码和边界框、场景元数据以及解释目标。我们使用LLaVA-1.5-7B、Qwen2.5-VL-7B和InternVL3.5-8B作为诊断基线而非所提方法。在包含1650张图像的合成测试集上,适配后的基线在标准留出场景上达到64.0%至67.8%的类别宏F1分数。对于LLaVA-1.5-7B,类别宏F1分数从标准划分上的64.0%下降到干预转移下的40.8%。这一差距表明,高分布内准确性部分反映了与渲染和反事实构建相关的线索。

英文摘要

Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language. We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba~3 to generate 5,500 images from 500 scene families. Each family contains one clean image and ten matched violations involving shadows, reflection, support, surface response, and occlusion. The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets. We use LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B as diagnostic baselines rather than proposed methods. On a 1,650-image synthetic test set, the adapted baselines reach 64.0--67.8\% category macro-F1 on standard held-out scenes. For LLaVA-1.5-7B, category macro-F1 falls from 64.0\% on the standard split to 40.8\% under intervention shift. This gap shows that high in-distribution accuracy partly reflects cues tied to rendering and counterfactual construction.

CommentsAccepted in NeurIPS 2026 the 1st Workshop on Physical World AI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑