arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REVEAL:用于长视频问答的基于 rubric(评分标准)的显式证据充分性验证智能体

REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering

Caijun Yan, Yang Zhou, Meixing Shi, Haoran Sun, Yichen Li, Yuxiang Cai, Yankai Jiang

arXiv 2608.08612首次发表:更新:

AI 中文总结

该研究针对长视频问答中现有方法割裂事件、忽略证据充分性的问题,提出REVEAL框架,通过自适应分块的结构化记忆与rubric引导的证据验证,实现更可靠的推理且性能优于现有SOTA方法。

AI 中文摘要

近年来,检索增强和记忆增强方法已成为长视频问答领域两种有前景的范式。然而,现有方法通常依赖于固定时长的刚性时间分块(例如10秒)和静态离线记忆库,这不仅会割裂连贯的连续事件,还无法在实时推理过程中自适应调整。此外,无论是使用多尺度摘要还是多模态知识图谱,当前方法都优先关注检索的相关性,却忽略了证据的充分性——即使关键的时间、因果或细粒度动作证据仍然缺失,它们往往在仅检索到语义相关线索后就停止回答。为解决这些挑战,我们提出了REVEAL,一个基于rubric(评分标准)的智能体框架。作为基础,我们引入了一种基于自适应视觉相似度的预处理流水线,将视觉上连贯的相邻帧分组为自然事件单元,以构建离线-在线视频记忆:在离线阶段捕获全局视频上下文,同时在在线阶段动态维护基于问题条件的记忆。基于这种结构化记忆,REVEAL利用自动构建的rubric库显式验证检索到的证据是否满足充分性标准,在验证失败时定位缺失的线索,并引导针对性的重新检索以获取补充信息。无需额外训练,REVEAL在大量实验中始终优于闭源和开源的最先进方法。这些结果表明,显式验证证据充分性而非仅停留在语义相关性,能够检索到现有方法遗漏的决定性线索,从而产生更可靠的长视频推理。

英文摘要

Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. However, existing methods typically rely on rigid, fixed-length temporal chunking (e.g., 10s) and static offline memory banks, which not only fragment coherent continuous events but also fail to adapt during real-time reasoning. Moreover, whether using multi-scale summaries or multimodal knowledge graphs, current approaches prioritize retrieval relevance while overlooking evidence sufficiency, often stopping to answer once only semantically relevant clues are retrieved, even when key temporal, causal, or fine-grained action evidence is still missing. To tackle these challenges, we propose REVEAL, a rubric-guided agent framework. As a foundation, we introduce an adaptive visual-similarity-based preprocessing pipeline that groups visually coherent adjacent frames into natural event units to construct an offline-online video memory---capturing global video context offline while dynamically maintaining question-conditioned memory online. Built upon this structured memory, REVEAL uses an automatically constructed rubric library to explicitly verify whether retrieved evidence satisfies sufficiency criteria, pinpoints missing clues upon verification failure, and directs targeted re-retrieval for complementary information. Without any extra training, REVEAL consistently outperforms both closed-source and open-source state-of-the-art methods across extensive experiments. These results show that explicitly verifying evidence sufficiency, rather than stopping at semantic relevance, retrieves the decisive clues that prior methods miss and yields more reliable long-video reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑