arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeoAnchor:通过潜在分解进行协作推理以实现3D空间理解

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

Hao Li, Han Fang, Zixin Pan, Xin Wei, Hongbo Sun, Jinglin Xu, Zhiyu Lin, Ye Yuan, Zhongjiang He, Yu Yu, Hao Sun

arXiv 2607.13454首次发表:更新:

发表机构

Shanghai Jiao Tong University; Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd; University of Science and Technology Beijing; The Hong Kong University of Science and Technology (Guangzhou)(上海交通大学; 星辰通用人工智能实验室,中国电信人工智能技术(北京)有限公司; 北京科技大学; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对从2D图像理解3D空间关系的挑战,提出GeoAnchor框架,通过分解3D空间信息为互补组件并结合协作训练策略,实现动态可解释推理,在复杂3D推理任务中优于现有技术。

AI 中文摘要

尽管多模态大语言模型取得了显著进展,但从2D图像理解3D空间关系仍是一项关键挑战。现有方法主要依赖符号文本标记,缺乏表示连续几何信息的保真度。近期方法虽用潜在表示增强推理,但单一潜在类型无法适应空间任务多样性。为此提出GeoAnchor,一个交错文本-潜在推理框架。它将3D空间信息分解为三个互补组件,在结构化空间中重组以构建局部证据并捕捉全局上下文,实现动态可解释推理。还引入协作训练策略。实验表明GeoAnchor优于现有技术,验证了其有效性和泛化能力。

英文摘要

Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing methods primarily rely on symbolic text tokens, which inherently lack the fidelity to represent continuous geometric information. While recent methods use latent representations to enhance reasoning, relying on a single latent type cannot adapt to the diversity of spatial tasks, leading to misalignment in complex geometric scenarios. To address these limitations, we propose GeoAnchor, an interleaved text-latent reasoning framework. GeoAnchor decomposes 3D spatial information into three complementary components: position latents for object grounding, direction latents for relational orientation, and geometry latents for scene structure. These components are recombined in a structured space to construct local evidence while capturing global context, enabling dynamic and interpretable reasoning. Furthermore, we introduce a collaborative training strategy that guides the model from local spatial perception to comprehensive 3D understanding. Extensive experiments on diverse and complex 3D reasoning tasks demonstrate that GeoAnchor outperforms the state of the art, validating its effectiveness and generalization capabilities.

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑