arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14138cs.CVcs.AI

SPARGen:通过原生多模态生成统一空间感知与推理

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi, Kewang Deng, Zukai Chen, Feifei Shao, Lei Yang, Quan Wang, Yawei Luo

AI总结:

SPARGen是统一多模态框架,将3D重建等空间任务转为指令条件生成任务,在单一框架内实现异构空间任务的竞争力性能,突破现有方法的知识迁移限制。

AI中文摘要:

从视觉观测中进行空间感知与推理,需要恢复几何结构、建立对应关系并理解空间关系。现有方法通常使用特定任务的架构或外部几何模块分别处理这些能力,限制了同一物理场景互补表示之间的知识迁移。我们提出SPARGen,这是一个统一的多模态框架,将3D重建、密集对应和空间推理转化为指令条件下的生成任务。SPARGen将紧凑的结构化输出和语言输出序列化为token序列,同时生成与图像对齐的密集几何场,使空间监督能够在原生多模态生成模型内共同塑造共享表示。在3D重建、对应和空间推理的基准测试实验表明,SPARGen在单一原生多模态生成框架内,在异构空间任务中实现了具有竞争力的性能。

英文摘要:

Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.

↑