AI 中文总结
针对自我中心助手处理复杂空间查询的挑战,提出空间接地对话记忆(SpaC-MEM),利用3D重建与分割压缩多模态历史,并构建Ego-SpaCR基准,在准确率与对象回忆上优于基线,验证了空间证据的关键作用。
AI 中文摘要
自我中心助手必须将用户所说的话与他们在长时间交互历史中所看到的内容联系起来。我们将这一挑战形式化为空间接地对话推理(SpaCR):跨场景、面向回忆和反事实的空间查询,这些查询将用户陈述的事实与几何证据相结合。随着历史增长,直接的视觉语言模型会产生高推理成本和上下文限制,而关键帧选择和检索可能会遗漏完整回忆所需的对象或证据。我们提出了空间接地对话记忆(SpaC-MEM),一种以对象为中心的工作记忆,利用3D重建和分割将对话信息接地于持久的物理对象。它压缩多模态历史,同时保留空间证据,并允许通过对话更新对象特定的事实。我们还引入了Ego-SpaCR,一个包含620个ScanNet视频会话的基准,并增加了95个面向任务的对话和3,100个评估查询。SpaC-MEM在评估方法中实现了最高的整体答案准确率,并提高了对象回忆,同时相比原生视频基线需要大幅更少的参考输入令牌。移除3D空间信息会显著降低性能,突显了同时保留空间和对话证据的重要性。
英文摘要
Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that combine user-stated facts with geometric evidence. Direct vision-language models incur high inference costs and context limits as histories grow, while keyframe selection and retrieval can omit objects or evidence needed for complete recall. We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. It compresses multimodal histories while preserving spatial evidence and allowing object-specific facts to be updated through dialogue. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines. Removing 3D spatial information substantially degrades performance, highlighting the importance of preserving spatial and conversational evidence together.