arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ICM-Bench:具备长期记忆的多模态智能体中的人物级身份推理

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

Shidu Ren, Yunze Liu, Xing Liu, Chi-Hao Wu, Enmin Zhou, Junxiao Shen

arXiv 2609.04438首次发表:更新:

发表机构

University of Toronto; Tsinghua University; Northeastern University; University of Bristol(多伦多大学; 清华大学; 东北大学; 布里斯托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ICM-Bench是首个评估多模态智能体长视频记忆身份推理的基准,含839个视频片段与1217个问题,实验显示Gemini 3.1 Pro在身份相关任务上表现仍有不足。

AI 中文摘要

长程多模态智能体不仅应记住发生了什么,还应记住谁参与其中。该能力依赖于将反复出现的人脸、语音、姓名、与人物相关的物品、事件及社会关系与随时间保持一致的身份关联起来。现有的长视频及多模态智能体基准测试衡量的是宽泛的记忆问答能力,但未将维持反复出现的人物身份并推理其跨时间关系的能力单独分离出来。我们推出ICM-Bench(以身份为中心的记忆基准),据我们所知,这是首个专门用于评估多模态智能体针对长视频记忆进行以身份为中心的推理的基准。该基准包含839个合成片段,时长总计141分钟,以及关于一本一年生活相册中6位反复出现的成年人的1217个开放式问题。一条主题可配置的流水线生成视频集合,并为每个问题关联其目标身份及可追溯的支撑证据。我们对比了直接字幕记忆基线、记忆增强智能体及图检索系统。Gemini 3.1 Pro取得了74.0%的最高总体准确率,但在需要长期身份档案的问题上,其得分降至60.3%。结果表明,当前系统能恢复许多事件级记忆,但在必须围绕稳定人物积累证据时,可靠性仍较低。

英文摘要

Long-horizon multimodal agents should remember not only what happened but also who participated. This capability depends on linking recurring faces, voices, names, person-associated objects, events, and social relations to consistent identities over time. Existing long-video and multimodal-agent benchmarks measure broad memory question answering, but they do not isolate the ability to maintain recurring person identities and reason over their cross-time relations. We introduce ICM-Bench (Identity-Centric Memory Benchmark), which, to the best of our knowledge, is the first benchmark specifically designed to evaluate identity-centric reasoning over long video memories in multimodal agents. The benchmark contains 839 synthetic clips spanning 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. A theme-configurable pipeline generates the video collection and associates each question with its target identities and traceable supporting evidence. We compare direct caption-memory baselines, memory-augmented agents, and graph-retrieval systems. Gemini 3.1 Pro achieves the highest overall accuracy of 74.0%, yet its score falls to 60.3% on questions that require long-term identity profiles. The results show that current systems recover many event-level memories but remain less reliable when evidence must be accumulated around a stable person.

Comments20 pages, 6 figures, and 7 tables. Code: https://github.com/Shidu-Ren/ICM-Bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑