arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeoLatent:基于路由优化的几何引导潜在结构用于3D推理

GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

Yakun Zhu, Yi Bin, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Duo Peng, Jingkuan Song, Heng Tao Shen

arXiv 2610.02091首次发表:更新:

发表机构

Tongji University; The Hong Kong Polytechnic University(同济大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GeoLatent通过CR-GEO和路由优化结构化几何潜在表示,提升3D空间推理,在SPAR-Bench和SPBench上分别达到73.0%和72.1%。

AI 中文摘要

尽管视觉-语言模型取得了进展,但从2D图像进行3D空间推理仍然具有挑战性。基于文本的方法用离散标记描述中间几何,限制了连续空间关系的保真度。连续潜在表示提供了更丰富的表示,但单一潜在类型不能明确分离跨空间任务所需的线索。分解的空间潜在表示通过分别表示位置、方向和全局几何,并在几何监督下解决这一问题。然而,几何表示仍可能坍缩为一个主导方向,且不受限制的注意力可能导致潜在表示在答案学习过程中未被充分利用。我们提出GeoLatent,结合公共-残差几何对齐(CR-GEO)与路由优化来结构化几何状态,同时促进潜在介导的答案学习。CR-GEO分离共享教师几何与残差教师几何;路由优化联合训练几何和语言,暂时通过潜在表示引导视觉答案学习,并在几何监督下恢复完整注意力。在受控比较中,CR-GEO将几何有效秩从1.00提升至3.87,而在瓶颈处阻断潜在读出会使128个固定问题上的方向准确率从89.1%降至25.8%。恢复后,差异化的几何表示和潜在介导的视觉路径仍与直接图像访问并行可用。GeoLatent在SPAR-Bench上达到73.0%,在SPBench上达到72.1%,在两个基准上均优于先前报道的方法。

英文摘要

Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.

Comments23 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑