arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CASE:面向高效长视频智能体的成本感知停止机制

CASE: Cost-Aware Stopping for Efficient Long-Video Agents

Yiming Du, Chenghao Liu, Zhiyuan Liu, Fangxing Zheng, Zhao Wang, Junnan Nie, Songfang Huang

arXiv 2610.05400首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CASE提出成本感知的序列停止框架,通过轻量回归器决定长视频智能体何时停止证据收集,在多个基准上显著节省令牌和运行时间,同时保持或提升准确率。

AI 中文摘要

长视频智能体能够主动收集与问题相关的证据,但它们通常留下一个核心决策悬而未决:智能体何时已经看得足够多从而可以回答?我们提出CASE,一个即插即用的终止框架,将该决策建模为策略条件化的序列停止问题。在每个因果检查点,CASE将累积证据的辅助多项选择评估与宿主智能体的执行状态相结合。从完整的原生轨迹中,我们构建了一个成本感知目标,该目标将立即回答与在同一搜索路径上稍后停止进行比较,同时考虑回答正确性和继续推理的全部成本。一个轻量级的Ridge回归器学习这一决策差距,并产生停止/继续决策。我们使用VideoSeek和AVP评估了三个视觉语言模型。在Video-MME上,端到端准确率平均变化+0.67个百分点,而CASE将模型令牌使用量减少了53.63%。相同的冻结策略随后零样本迁移到LongVideoBench和MLVU,端到端准确率变化分别为+3.38和+4.58个百分点,同时分别节省了58.78%和51.28%的模型令牌。在所有智能体-模型-基准组合中,CASE在所选操作点上达到了所比较的停止方法中最高的准确率-效率帕累托前沿覆盖率(83.3%)。在线执行保持了这种有利的准确率-效率权衡,并额外将实测运行时间平均减少了54.1%。CASE为长视频推理智能体提供了一个即插即用的终止框架,使它们能够决定何时进一步获取证据不再值得。

英文摘要

Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned sequential stopping. At each causal checkpoint, CASE combines an auxiliary multiple-choice assessment of accumulated evidence with the host agent's execution state. From complete native trajectories, we construct a cost-aware target that compares answering now with stopping later along the same search path, accounting jointly for answer correctness and the full cost of continued reasoning. A lightweight Ridge regressor learns this decision gap and produces STOP/CONTINUE decisions. We evaluate three vision-language models with VideoSeek and AVP. On Video-MME, end-to-end accuracy changes by +0.67 percentage points on average while CASE reduces model-token use by 53.63%. The same frozen policies then transfer zero-shot to LongVideoBench and MLVU, with end-to-end accuracy changes of +3.38 and +4.58 points while saving 58.78% and 51.28% of model tokens, respectively. Across all agent-model-benchmark combinations, CASE attains the highest accuracy-efficiency Pareto-frontier coverage among the compared stopping methods (83.3%) at the selected operating points. Online execution preserves this favorable accuracy-efficiency trade-off and additionally reduces measured runtime by 54.1% on average. CASE provides a plug-in termination framework for long-video reasoning agents, enabling them to decide when further evidence acquisition is no longer worthwhile.

Comments40 pages, including references and appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑