发表机构
The Hong Kong University of Science and Technology; Xiaohongshu Inc.; The University of Hong Kong; The Chinese University of Hong Kong(香港科技大学; 小红书公司; 香港大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对当前智能体流媒体视频理解评估的缺陷,提出StreamArena基准,设计StreamMind架构解决连续交互与长视界理解的张力,在四项能力上优于现有基线并降低延迟。
AI 中文摘要
在连续的真实环境中部署自主多模态智能体,要求其能够接收无限的音视频流并维持小时级别的记忆。然而,当前的评估主要依赖于简短的视频片段和多项选择题形式,这种设计使得仅处理最后四帧的极简基线模型就能达到或超过复杂的流媒体模型,且答案选项也会暴露语言捷径。我们推出StreamArena,这是一个用于小时级交互式流媒体视频理解的基准,包含243个平均时长88.8分钟的完整视频,以及3646个经过严格标注的开放式问答对,用于评估实时感知、历史回顾、主动交互和多模态工具利用能力。对不同系统的评估揭示了连续交互与长视界多模态理解之间的张力:仅保留最近帧的方法无法恢复远处事件,将过去观测转换为文本的方法会丢失视觉证据,反复压缩视觉记忆的方法难以随时间保留细粒度细节。我们通过StreamMind解决了这一张力,这是一个两层架构,将延迟关键型交互和主动监控分配给独立调度的前端工作节点,而后端工作节点异步构建持久多模态记忆并执行历史回忆和外部搜索。StreamMind在所有四项能力上均优于现有的流媒体基线模型,并通过复用持久状态降低了查询到答案的延迟。
英文摘要
Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design allows minimal baselines that process only the last four frames to match or surpass complex streaming models, while answer options also expose language shortcuts. We introduce StreamArena, a benchmark for hour-scale, interactive streaming video understanding. StreamArena contains 243 full-length videos averaging 88.8 minutes and 3,646 rigorously annotated, open-ended question-answer pairs that evaluate real-time perception, historical retrospection, proactive interaction, and multimodal tool utilization. Evaluation across diverse systems exposes a tension between continuous interaction and long-horizon multimodal comprehension. Methods that retain only recent frames cannot recover distant events, methods that convert past observations into text lose visual evidence, and methods that repeatedly compress visual memory struggle to preserve fine-grained details over time. We address this tension with StreamMind, a two-tier architecture that assigns latency-critical interaction and proactive monitoring to independently scheduled frontend workers, while backend workers asynchronously construct persistent multimodal memory and perform historical recall and external search. StreamMind outperforms existing streaming baselines across all four capabilities and reduces query-to-answer latency by reusing persistent state.