arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Slot3R:用于流式三维重建的组关联空间记忆

Slot3R: Set-Associative Spatial Memory for Streaming 3D Reconstruction

Xiyuan Zhang, Yanming Yang, Kaiyuan Xu, Ruibo Li, Chi Zhang

arXiv 2610.12282首次发表:更新:

发表机构

Westlake AGI Lab(西湖AGI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Slot3R是一种无需训练的组关联空间记忆改造,冻结Point3R骨干网络,在300-500帧时显著降低点云精度误差与绝对轨迹误差,且内存效率优于Point3R和InfiniteVGG。

AI 中文摘要

流式三维重建需要在线处理不断扩展的场景,同时保留每一帧的证据。空间记忆是合适的选择,因为它按重建的三维位置组织历史信息。然而Point3R将空间邻近性既用于将新观测与现有记忆条目关联,又用于决定是否融合,将共位置与状态身份混为一谈。由于指针汇总图像块,邻近指针可能编码不同的表面、视点或可见性条件;在后续帧消除歧义前对其求平均会破坏互补证据。我们认为位置应决定地址,而非观测是否必须合并。Slot3R将此原则实现为无需训练的组关联改造,在冻结预训练Point3R骨干网络的同时,允许多个状态在同一地址共存。有界稀疏读出进一步将持久存储与每帧解码器访问解耦。在300-500个采样帧时,Slot3R在7Scenes上将Point3R的点云精度误差(Acc)降低57.1%-63.1%,在NeuralRGBD上降低64.0%-72.0%;在所有三个姿态基准上降低Sim(3)对齐的绝对轨迹误差(ATE),且在视频深度估计上保持竞争力。在相同协议下,其在600至1000个采样帧的所有评估设置中以约19 FPS运行,而Point3R和InfiniteVGG在800帧及以上时内存耗尽。

英文摘要

Streaming 3D reconstruction must preserve evidence from each frame while processing an expanding scene online. Spatial memory is a natural fit because it organizes history by reconstructed 3D location. Yet Point3R uses spatial proximity both to associate a new observation with an existing memory entry and to decide whether to fuse it, conflating co-location with state identity. Because pointers summarize image patches, nearby pointers may encode distinct surfaces, viewpoints, or visibility conditions; averaging them can destroy complementary evidence before later frames disambiguate it. We argue that location should determine address, not whether observations must merge. Slot3R realizes this principle as a training-free, set-associative retrofit that lets multiple states coexist at a shared address while keeping the pretrained Point3R backbone frozen. A bounded sparse readout further decouples persistent storage from per-frame decoder access. At 300-500 sampled frames, Slot3R reduces Point3R's point-cloud accuracy error (Acc) by 57.1%-63.1% on 7Scenes and 64.0%-72.0% on NeuralRGBD, lowers Sim(3)-aligned absolute trajectory error (ATE) on all three pose benchmarks, and remains competitive on video-depth estimation. It completes all evaluated settings from 600 to 1000 sampled frames at about 19 FPS under the same protocol, whereas Point3R and InfiniteVGGT run out of memory at 800 frames and beyond.

CommentsProject Page: https://ashleyxyz.github.io/Slot-3R/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑