arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CADER:用于长视频理解的置信度感知动态证据推理

CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

Jinlong Yang, Wenhao Zhang, Kuanwei Lin, Sijie Cheng

arXiv 2607.24582首次发表:更新:

AI 中文总结

针对长视频理解中现有系统推理过程不考虑难度的问题,提出CADER框架。它先全局推理估计置信度让高置信度示例提前退出,对不确定示例激活第二阶段工具增强循环定位证据,实验证明其能提升长视频推理能力并提供实用推理路线。

AI 中文摘要

长视频理解越来越依赖大型视觉语言模型和工具增强推理,但大多数系统对每个示例都采用相同的推理过程,而不考虑难度。这种统一策略对简单问题调用了不必要的工具辅助处理,在难题需要细粒度时间证据时提供的控制有限。我们提出了CADER(置信度感知动态证据推理),这是一个用于自适应和可靠长视频推理的无需训练的框架。CADER首先对均匀采样的帧进行全局推理,并用对数边际信号估计答案置信度,使高置信度示例能够提前退出。对于不确定的示例,CADER激活第二阶段工具增强循环,该循环结合了时间裁剪、轻量级语义验证和相关性引导重采样,以逐步定位与问题相关的证据。实验表明,CADER提高了长视频推理能力,且高置信度样本可绕过第二阶段。此外,当应用于仅通过无工具思维链监督训练的主干时,CADER与专门的工具增强框架相比具有竞争力,为自适应长视频推理提供了一条实用的推理时间路线。

英文摘要

Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example regardless of difficulty. This uniform strategy invokes unnecessary tool-assisted processing for easy questions and provides limited control when difficult questions require fine-grained temporal evidence. We propose CADER (Confidence-Aware Dynamic Evidence Reasoning), a training-free framework for adaptive and reliable long-video reasoning. CADER first performs global reasoning over uniformly sampled frames and estimates answer confidence with a logit-margin signal, allowing high-confidence examples to exit early. For uncertain examples, CADER activates a second-stage tool-augmented loop that combines temporal cropping, lightweight semantic verification, and Relevance-Guided Resampling to progressively localize question-relevant evidence. This design treats tool use as a sample-level decision: a single global pass handles easy cases, while additional reasoning is reserved for examples where uncertainty suggests that more evidence is needed. Experiments on multiple VideoQA benchmarks show that CADER improves long-video reasoning while bypassing Stage~2 for high-confidence samples. Moreover, when applied to a backbone trained only with tool-free chain-of-thought supervision, CADER achieves competitive performance against specialized tool-augmented frameworks, suggesting a practical inference-time route for adaptive long-video reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑