arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VAMR:面向高效长视频理解的多问题智能体推理

VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding

Runquan Gui, Hanzhu Chen, Zehao Wang, Hanxin Zhu, Xin Li, Zhibo Chen

arXiv 2610.11171首次发表:更新:

发表机构

University of Science and Technology of China; Tencent(中国科学技术大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出VAMR多问题视频智能体,通过共享工具轨迹处理长视频多问题,在三个基准数据集上实现最高准确率与最少推理轮次,性能优于VideoARM。

AI 中文摘要

长视频理解通常涉及对同一视频不同方面的多个问题,但现有视频智能体通常通过独立的工具使用轨迹处理每个问题,这会反复重启视频探索和记忆构建,错失联合获取证据、逐步构建支持全部问题集的共享理解的机会。我们提出VAMR(Video Agent for Multi-Question Reasoning,面向多问题推理的视频智能体),通过一条共享的工具使用轨迹协调视频的所有问题。每一轮中,持久策略模型可调用工具处理一个或多个未解决问题,并为有足够证据的问题提交答案;问题条件视觉感知可在一次调用中为多个问题检索细粒度线索,而分层多问题记忆将可重用上下文整合为共享视频故事,并为单个问题保留独立证据。在通过监督微调初始化该交互协议后,我们提出问题视野策略优化(qhpo),用于优化问题在不同轮次推进和完成的共享轨迹:具体而言,问题级评论员估计每个活跃问题的价值,同时轮次对齐将每个问题优势映射到直接服务于它的轮次,之后聚合对齐后的优势以优化共享演员。在LVBench、Video-Holmes和LongVideoBench数据集上,VAMR在迭代方法中实现了最高的整体准确率和最少的推理轮次;在LVBench上,其准确率达62.1%,超过VideoARM 4.3个百分点,同时推理轮次和处理帧分别减少85.9%和61.4%。

英文摘要

Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficient evidence. Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions. After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds. Specifically, a question-level critic estimates the value of each active question, while round alignment maps each question advantage to the rounds that directly serve it before the aligned advantages are aggregated to optimize the shared actor. Across LVBench, Video-Holmes, and LongVideoBench, VAMR achieves the highest accuracy overall and the fewest reasoning rounds among iterative methods. On LVBench, it reaches 62.1\% accuracy, exceeding VideoARM by \textbf{4.3} points while reducing reasoning rounds and processed frames by \textbf{85.9\%} and \textbf{61.4\%}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑