CapMem:基于字幕的自我中心视频情景记忆基准
CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
- University College London(伦敦大学学院)
- Imperial College London(帝国理工学院)
- Durham University(杜伦大学)
- Shanghai Jiao Tong University(上海交通大学)
- Queen Mary University of London(伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出CapMem基准,验证在长自我中心视频中,文本字幕可作为可重用的情景记忆,通过字幕问答和检索验证框架显著提升推理准确率。
AI中文摘要:
可穿戴助手需要对自我中心视频具备情景记忆能力,然而当前的视觉语言模型面临有限的帧预算、不断增长的视觉标记成本以及长上下文检索失败等问题。在这些实际约束下,我们研究文本字幕是否可以作为可重用的情景记忆。我们定义了情景记忆视频字幕问答任务,并引入了CapMem,这是一个人工标注的基准,包含75个视频,总时长33.7小时,以及跨越16个场景的1,000道多项选择题。在长视频(超过20分钟)上,使用30秒和60秒字幕窗口的全覆盖字幕问答分别优于直接视频问答,对于10/12和8/12的模型。在相同的视频子集上,对六个Qwen模型进行的匹配帧控制保留了平均准确率提升,分别为3.22和2.55个百分点。我们的字幕引导的检索与验证框架进一步将准确率提升了最多5.3个百分点。这些结果支持了字幕记忆在长自我中心视频情景推理中的有效性。
英文摘要:
Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-context retrieval failures. Under these practical constraints, we study whether textual captions can serve as reusable episodic memory. We define the Episodic Memory Video Caption QA task and introduce CapMem, a human-annotated benchmark with 75 videos totaling 33.7 hours, and 1,000 multiple-choice questions across 16 scenarios. On long videos (>20 min), full-coverage CaptionQA with 30s and 60s caption windows outperforms direct VideoQA for 10/12 and 8/12 models, respectively. On the same video subset, a matched-frame control across six Qwen models retains mean accuracy gains of 3.22 and 2.55 points, respectively. Our caption-guided retrieve-and-verify harness further improves accuracy by up to 5.3 points. These results support the effectiveness of caption memory for episodic reasoning over long egocentric video.