发表机构
Institute for AI Industry Research, Tsinghua University; Peking University(清华大学人工智能产业研究院; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MemEvo利用LLM驱动的自动研究框架,通过领域特定语言搜索可执行记忆程序,无需训练即可发现高效的有界流式视频记忆机制,在多个基准上验证了其性能与效率。
AI 中文摘要
查询无关的流式视频理解要求视觉语言模型在未来的查询已知之前,将无限增长的视觉流持续压缩为有界记忆。其性能关键取决于记忆机制——保留哪些观察,如何表示和整合它们,以及当查询最终到达时检索哪些信息。我们不是手工设计单一的记忆架构,而是将记忆设计表述为对可执行记忆程序的搜索问题。我们引入了一种轻量级的领域特定语言,通过表示、准入、保留、整合、预算和检索等结构化原语来表达记忆机制,同时强制执行因果和有界记忆约束。尽管是结构化的,所推导的程序空间仍然很大,包含异构的、条件依赖的设计选择,其效果只能通过下游执行来评估。因此,我们提出了MemEvo,一个由LLM驱动的自动研究框架,利用预训练的LLM作为语义感知的提议模型,基于累积的实验反馈迭代生成和细化候选记忆程序。在运行时,一个确定性的评估流水线验证并评估每个候选,而底层的视觉语言模型在整个发现过程中保持冻结。我们最终产生了一个无需训练的有界记忆机制。在StreamingBench和OVO-Bench上的大量实验展示了强大的流式视频理解性能,以及显著的情境和推理效率。
英文摘要
Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses memory mechanisms through structured primitives for representation, admission, retention, consolidation, budgeting, and retrieval, while enforcing causal and bounded-memory constraints. Although structured, the derived program space remains large and contains heterogeneous, conditionally dependent design choices whose effects can only be assessed via downstream execution. We therefore propose MemEvo, an LLM-driven auto-research framework that uses pretrained LLM as a semantics-aware proposal model to iteratively generate and refine candidate memory programs based on accumulated experimental feedback. At runtime, a deterministic evaluation pipeline validates and evaluates each candidate, while the underlying vision-language model remains frozen throughout discovery. We finally produce a training-free, bounded-memory mechanism. Extensive experiments on StreamingBench and OVO-Bench demonstrate strong streaming video understanding performance together with substantial context and inference efficiency.