arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19059cs.ROcs.CV

LT-Mem:面向终身场景理解的感知时间动态的时空记忆

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

Yumin Lee, Hyoseok Ju, Giseop Kim

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对机器人长期场景理解的时间遗忘问题,提出LT-Mem时空记忆框架,结合多会话SLAM与Tri-Memory结构,在LT-VQA数据集上性能优于基线且token消耗更少。

中文摘要 AI 辅助

长期在动态环境中运行的机器人需要能在多次回访中持续存在的物体级理解。现有系统要么覆盖历史以维持最新地图,要么存储语义快照但缺乏跨会话一致的物体身份,导致时间遗忘症:系统性丢失物体历史,无法回答如“绿色椅子在所有会话中去过哪里?”这类查询。我们提出LT-Mem,一种感知时间动态的记忆演化框架,它将空间对齐的实例级3D感知与时间动态条件推理统一起来。首先,多会话SLAM主干提供跨会话空间对齐的每个物体观测。其次,推理层控制物体记忆的演化:确定性证据评分保留跨会话身份,感知时间动态的策略根据每个物体的动态在覆盖、保留和多假设行动中选择。第三,生成的Tri-Memory结构(Live、Delta、Meta)同时保留当前状态和事件历史,支持以物体为中心的纵向推理。我们进一步引入LT-VQA,一个包含多会话记录、持久身份注释和时间问答对的数据集与评估套件。实验表明,LT-Mem在所有指标上均持续优于基线,同时消耗的token数量少一个数量级, ablation实验证实性能提升源于结构化记忆架构而非LLM能力。

英文摘要

Long-term robot operation in evolving environments requires object-level understanding that persists across repeated revisits. Existing systems either overwrite history to maintain an up-to-date map or store semantic snapshots without consistent cross-session object identity, resulting in temporal amnesia: the systematic loss of object history that prevents answering queries such as "Where has the green chair been across all sessions?" We propose LT-Mem, a volatility-aware memory evolution framework that unifies spatially aligned instance-level 3D perception with volatility-conditioned temporal reasoning. First, a multi-session SLAM backbone provides spatially aligned per-object observations across sessions. Second, a reasoning layer governs how object memory evolves: deterministic evidence scoring preserves cross-session identity, and a volatility-aware policy selects among overwrite, hold, and multi-hypothesis actions based on each object's dynamics. Third, the resulting Tri-Memory structure (Live, Delta, Meta) preserves both current states and event histories, enabling longitudinal object-centric reasoning. We further introduce LT-VQA, a dataset and evaluation suite comprising multi-session recordings, persistent identity annotations, and temporal QA pairs. Experiments show that LT-Mem consistently outperforms baselines across all metrics while consuming an order of magnitude fewer tokens, and ablations confirm that gains are driven by the structured memory architecture rather than LLM capacity.

发表机构

  • DGIST(大邱庆北科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑