发表机构
FAIR, Meta; Reality Labs, Meta(公平研究院,元公司; 现实实验室,元公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对可穿戴设备连续第一人称记录下AI助手的情景记忆问题,引入S-EMBER基准测试,通过视频及问答对模拟真实交互,发现前沿模型存在定位矛盾,为可穿戴AI代理情景记忆发展奠定基础。
AI 中文摘要
随着可穿戴设备实现连续第一人称记录,AI助手必须跨长时间跨度进行推理以回忆过去经历,即情景记忆。当前基准测试常依赖离线评估,无法模拟可穿戴智能的流式现实。我们引入S-EMBER,一个包含3141个视频共388小时有机活动的大规模基准测试,形式化了基于视觉事件触发的因果主动回忆的流式情景检索,揭示了定位矛盾,为下一代可穿戴AI代理开发可靠情景记忆奠定硬件真实基础。
英文摘要
As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory. Current benchmarks often rely on offline evaluation with access to entire video files, failing to simulate the streaming reality of wearable intelligence. We introduce S-EMBER (Streaming Egocentric Memory Benchmark for Episodic Retrieval), a large-scale benchmark comprising 3,141 videos totaling 388 hours of organic activity captured via Ray-Ban Meta smart glasses. S-EMBER formalizes grounded streaming episodic retrieval, a paradigm shift from global offline search to causal, active recall triggered by visual events in a continuous stream. We provide 9,448 QA pairs requiring manual visual proof through precise temporal localization and supporting flexible response lengths to simulate natural human-AI interaction. Our extensive benchmarking of frontier models reveals a grounded recall gap: models answer and localize with moderate competence in isolation, yet fall furthest short of human performance when both must hold for the same query, the strongest reaching less than half the human rate. S-EMBER establishes a hardware-authentic foundation for developing grounded, reliable episodic memory in the next generation of wearable AI agents.