发表机构
Nanjing University; University of British Columbia; Zhejiang Sci-Tech University; Huawei Technologies Co., Ltd.; Tianjin University(南京大学; 不列颠哥伦比亚大学; 浙江理工大学; 华为技术有限公司; 天津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出语言轨迹编码(LTE),构建空间记忆基准(SMB),在SMB和Ego4D上实现更优的长时程空间记忆与检索性能,压缩效率高且查询延迟低。
AI 中文摘要
执行长时程任务的具身智能体需要一种记忆表示,其中动态物体的状态转换在数小时至数天的观测时程内可通过自然语言查询。现有系统要么丢弃细粒度运动(片段级视频-语言嵌入),仅将其保留为原始坐标(几何SLAM),要么围绕即时任务上下文组织(智能体工作记忆),均未为智能体提供可通过语言查询其自身状态转换的逐物体时间线。本文的核心贡献是语言轨迹编码(Linguistic Trajectory Encoding, LTE),它通过结合自然语言描述、稀疏空间锚点和视觉锚点的混合表示来压缩动态物体运动历史。LTE根据运动复杂度调整压缩方式:将无可靠观测的时间段锚定到最后可见位置,同时用几何路点和语言描述表示运动以保留准确性。为评估其在长时程的能力,本文从EgoLife多天录制数据中构建了空间记忆基准(Spatial Memory Benchmark, SMB),针对现有基准缺失的能力:语义轨迹检索和长时程物体检索。在SMB上,基于LTE的系统在语义轨迹检索中达到45.3%的成功率,在长时程物体检索中达到48.7%,优于结构化记忆和VLM基准(最佳先前结果分别为31.9%和34.4%)。LTE在24小时视频上实现了8.7倍至26.1倍的轨迹压缩,查询延迟低于1秒。在Ego4D自然语言查询上,该系统达到28.75%的R@1和55.10%的R@5,比EgoVLPv2提升15.80和31.30个百分点。
英文摘要
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (clip-level video-language embeddings), keep it only as raw coordinates (geometric SLAM), or organise it around immediate task context (agent working memories). None of them gives the agent a per-object timeline whose state transitions are themselves queryable in language. Our key contribution is \textbf{Linguistic Trajectory Encoding} (LTE), which compresses dynamic object motion histories via a hybrid representation combining natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location, while representing motion with geometric waypoints and linguistic descriptions to preserve accuracy. To evaluate these capabilities across extended time horizons, we construct the \textbf{Spatial Memory Benchmark} (SMB) from EgoLife multi-day recordings, targeting capabilities absent in existing benchmarks: semantic trajectory retrieval and long-horizon object retrieval. On SMB, the LTE-based system achieves $45.3\%$ success in semantic trajectory retrieval and $48.7\%$ in long-horizon object retrieval, outperforming structured-memory and VLM baselines (best prior: $31.9\%$ and $34.4\%$). LTE achieves trajectory compression by factors of $8.7\times$ to $26.1\times$ with sub-second query latency on $24$\,h video. On Ego4D natural-language queries, the system reaches $28.75\%$ / $55.10\%$ R@1/R@5, $+15.80$ / $+31.30$ pts over EgoVLPv2.
CommentsAccepted at NeurIPS 2026. Project page: https://sealical.github.io/st-mem/