arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37559cs.CV

APM-Bench:面向第一人称流式视频助手的跨会话持久记忆基准

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

Jianguo Huang, Jinming Liu, Qiyao Wang, Liang Xu, Jianhang Li, Zhimian Wen, Mingda Li, Shule Lu, Zhicheng Wang, Yuhan Guo, Xin Jin, Wenjun Zeng

首次发表
浏览论文内容

中文总结 AI 辅助

APM-Bench提出多会话生活轨迹基准,评估流式视频助手在间歇交互中的持久记忆,揭示效用-延迟-存储权衡,推动实用记忆系统开发。

中文摘要 AI 辅助

为了作为现实世界中的个人助手,流式视频模型需要具备持久记忆能力,以保留过去的经验供后续使用。然而,现有的流式基准和方法通常聚焦于单个连续视频或短视频片段,忽视了现实世界中的交互往往是间歇性的,并且需要记忆在中断期间持续存在。为填补这一空白,我们提出了APM-Bench,将现实世界的流式交互重新构建为多会话生活轨迹。它包含549个会话、104条轨迹和2,719个候选问题,涵盖客观题和开放性问题。每个会话是一个带有细粒度标注的视频,同一轨迹内的会话围绕相关活动展开。模型随后利用持久记忆回答关于过去会话的问题,并在保持实时交互的同时提供主动响应。这带来了挑战:持久记忆必须可存储、有选择性地保留信息、在适当时机注入,并保持高效。此外,有限的存储可能导致所需证据不可用,因此助手应能识别证据缺失。因此,我们在不同记忆协议和多种专门化流式记忆系统下系统评估了通用视频模型,并测试模型是否承认证据不足。我们的评估揭示了清晰的效用-延迟-存储权衡:现有方法仍难以在跨会话中同时实现可靠的长期回忆、低开销和有效的主动辅助。APM-Bench为在现实流式条件下开发和比较持久记忆系统提供了全面测试平台。我们希望它鼓励未来工作联合考虑效用、延迟和存储,以构建更实用的现实世界流式助手持久记忆。

英文摘要

To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Eastern Institute of Technology, Ningbo(宁波东方理工大学)
  • Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
  • Zhongguancun Academy(中关村学院)
  • Dalian University of Technology(大连理工大学)
  • Beihang University(北京航空航天大学)
  • Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑