arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

看见部分,推理世界:部分观测下的视觉推理

Seeing Parts, Reasoning about Worlds: Visual Inference under Partial Observation

Wei Wang, Wenqiao Zhang, Yutong Lin, Jun Xiao, Yueting Zhuang

arXiv 2609.32456首次发表:更新:

AI 中文总结

针对部分观测下的世界建模,提出WorldScope数据集和WorldFlow方法,通过可能世界语义和证据格推理,在5,000问测试集上达64.34%准确率,较基线提升24.88个百分点。

AI 中文摘要

部分观测下的世界建模需要推理与有限视觉证据相容的完整世界。被遮挡的物体和未观察到的区域可能留下多种可能的世界状态;额外的视角可以排除其他可能性,并强化观测所支持的结论。我们引入WorldScope,通过可能世界语义、基于证据的数据和学习到的视觉表示来研究这一过程。WorldScope-1.2M提供了120万个英文问答对,涵盖八个世界属性、十种任务界面和三种观测协议。其答案编码了已确认的事实、有支持的界限和未解决的可能性。补充监督包括4,800个经过认证的反世界组,这些组具有相同的基础观测但不同的隐藏物体配置和查询答案。这些组提供了模糊性的物理见证,以及用于世界兼容性和视角诱导排除的训练专用标签。我们提出WorldFlow,它将跨视角实体证据和支持表面覆盖组合成一个图像子集证据格。反世界兼容性和转换目标训练子集表示,以反映新观测如何约束可能世界。一个共享的答案生成器使用这些表示来预测最强支持的结论。WorldScope-Bench评估当视角被选择、组合、移除或排序时的声明判断和依赖证据的结论。在其5,000个问题的测试集上,WorldFlow达到了64.34%的精确准确率,比仅在问答上训练的相同骨干网络提高了24.88个百分点。在来自结构不相交场景的3,000个问题上,它保持了50.43%的准确率。

英文摘要

World modeling under partial observation requires reasoning about the complete worlds that remain compatible with limited visual evidence. Occluded objects and unseen regions can leave several world states possible; additional views can exclude alternatives and strengthen the conclusions supported by the observations. We introduce WorldScope to study this process through possible-world semantics, evidence-grounded data, and learned visual representations. WorldScope-1.2M provides 1.2 million English question-answer pairs spanning eight world properties, ten task interfaces, and three observation protocols. Its answers encode confirmed facts, supported bounds, and unresolved possibilities. Complementary supervision comprises 4,800 certified counterworld groups with equivalent base observations and different hidden object configurations and query answers. These groups provide physical witnesses of ambiguity and training-only labels for world compatibility and view-induced exclusions. We propose WorldFlow, which composes cross-view entity evidence and support-surface coverage into an image-subset evidence lattice. Counterworld compatibility and transition objectives train subset representations to reflect how new observations constrain possible worlds. A shared answer generator uses these representations to predict the strongest supported conclusion. WorldScope-Bench evaluates claim judgments and evidence-dependent conclusions as views are selected, combined, removed, or ordered. On its 5,000-question test set, WorldFlow reaches 64.34% exact accuracy, improving over the same backbone trained on QA alone by 24.88 percentage points. It retains 50.43% accuracy on the 3,000 questions from structure-disjoint scenes.

Comments24 pages, 5 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑