EgoExoMem: 跨视角记忆推理 over 同步的自身视角和外部视角视频
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
浏览论文内容
中文总结 AI 辅助
本文提出EgoExoMem,首个跨视角记忆推理基准,通过同步自身视角和外部视角视频进行跨视角记忆推理,利用E$^2$-Select方法实现高效的帧选择,实验表明自身和外部视角提供互补的记忆线索,但现有模型在基准测试中表现有限。
中文摘要 AI 辅助
自身视角记忆在具身智能中被广泛应用,但可能不足以进行全面的空间-时间推理。受人类从现场和观察者视角回忆的启发,我们引入EgoExoMem,首个跨视角记忆推理基准,包含2600个高质量MCQs,覆盖八个时间、空间和跨视角QA类型。为支持双视角检索,我们提出E$^2$-Select,一种无需训练的帧选择方法,结合基于相关性的预算分配与每视角k-DPP采样,以处理视角不对称性和跨视角时间一致性。实验表明,自身和外部视角提供互补的记忆线索,而现有MLLMs仍远未解决该基准:最佳模型仅达到55.3%。E$^2$-Select在帧选择和RAG基于的记忆基线中达到最先进的58.2%。进一步分析揭示了问题框架和答案定位之间的系统性视角偏好冲突,突显了跨视角记忆推理的新颖性和挑战性。
英文摘要
Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and observer perspectives, we introduce EgoExoMem, the first benchmark for cross-view memory reasoning over synchronized egocentric and exocentric videos. EgoExoMem contains $2.6K$ high-quality MCQs across eight temporal, spatial, and cross-view QA types. To support dual-view retrieval, we propose E$^2$-Select, a training-free frame selection method for synchronized ego-exo videos. It combines relevance-based budget allocation with per-view k-DPP sampling to handle view asymmetry and cross-view temporal consistency. Experiments show that ego and exo views provide complementary memory cues, while existing MLLMs remain far from solving the benchmark: the best model reaches only $55.3\%$. E$^2$-Select achieves state-of-the-art performance of $58.2\%$ over frame-selection and RAG-based memory baselines. Further analysis reveals systematic view-preference conflicts between question framing and answer grounding, underscoring the novelty and challenge of cross-view memory reasoning.
发表机构
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
- ETH Zurich(苏黎世联邦理工学院)
- University of Oxford(牛津大学)
- Hunan University(湖南大学)
机构由 AI 辅助整理,请以论文原文为准。