发表机构
Rightly Robotics(Rightly机器人公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对开放式视频流存储问题,提出面向实体的多媒体存储系统ReflectWorld-MM,由感知前端、分层长期记忆和实际应用实现三部分构成,在六个相关基准测试中取得最佳准确率,优于其他模型。
AI 中文摘要
构建能够持续观察世界、记住所见并基于积累经验进行推理的助手是一个长期目标,近期配备视频流长时记忆的多模态智能体备受关注。现有系统存在局限,本文提出ReflectWorld-MM,一种面向实体的开放式视频流多媒体存储系统。它由感知前端、分层长期记忆和实际应用实现三部分组成。在六个长视频和终身记忆基准测试中,ReflectWorld-MM在所有六项测试中均取得最佳准确率,优于强大的记忆智能体和前沿模型。
英文摘要
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest. Unfortunately, existing systems either keep their memory inside the model context or in a flat feature store, and organize it around frames rather than around the persistent entities a stream is really about, which confines them to bounded videos and weakens their ability to track who and what reappears over time. In this paper, we propose ReflectWorld-MM, an entity-oriented multimodal memory system for open-ended video streams. It consists of three parts. The first is a perception front-end that turns an audiovisual stream into entity-resolved observations under a bounded short-term memory. The second is a hierarchical long-term memory, grounded in human memory theory, that couples a multi-scale episodic memory, an evolving entity-centric semantic memory, and a procedural memory. The third is a complete realization, built for real-world operation, that ingests arbitrary streams and plugs into off-the-shelf assistants. Across six long-video and lifelong-memory benchmarks, ReflectWorld-MM achieves the best accuracy on all six, outperforming strong memory agents and a frontier model.