发表机构
Texas A&M University; Hamad Bin Khalifa University(德克萨斯农工大学; 哈马德·本·哈利法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出成本感知的配对协议,联合评估准确性与成本,揭示动态工具合成在视频问答中的效率与代价。
AI 中文摘要
智能视频问答(VideoQA)系统在推理过程中调用工具,但其工具库是固定的,因此重复的过程每次都要从原语重新构建。合成复合工具可以消除这一开销,但这种扩展是否有帮助很难评估:标准指标最终答案准确性忽略了推理成本,因此无法揭示系统如何转移成本。我们提出了一种成本感知的配对协议,用于审计工具增强的视频代理。该协议为每个问题配对两个完整系统在相同输入上的表现,并联合报告它们在准确性和成本上的净差异。对于每个问题,它将配对结果分为六组,由联合正确性和可见工具调用的变化定义,将保持准确性的效率提升与有害回归区分开。显著性通过McNemar检验和配对自助置信区间报告。我们在Dynamic-SAGE上实例化该协议,这是一个智能VideoQA框架,它合成、验证并持久注册可执行的复合工具以便在未见问题上重用,并在SAGE-Bench上对SAGE基线进行评估。审计揭示了一个标量准确性比较会遗漏的多轴概况:Dynamic-SAGE将准确性提高了7.5个百分点(p < 0.001),并将推理轮次和可见工具调用减少了约28%,同时转移而非减少了推理成本,因为令牌使用量增加了34%,成本增加了26%。在视觉和开放式问题上收益最大,在语言和多模态问题上中性,残余失败集中在困难的开放式问题上,此时管道工作最多。通过联合测量准确性和成本,该协议显示了管道级差异在何处可靠,何处不可靠。代码可在以下网址获取:https://this https URL。
英文摘要
Agentic Video Question Answering (VideoQA) systems produce answers through adaptive reasoning and tool-use trajectories, yet standard practice evaluates each system once and estimates uncertainty only across questions. This leaves a basic question untested: would the measured method effect survive if the evaluation were run again? We show that it need not. Using Static-SAGE and Dynamic-SAGE as a controlled case study, we repeat the paired comparison twice on identical SAGE-Bench question-video pairs, holding configuration, tool library, and scoring protocol fixed. In the first execution, Dynamic-SAGE outperforms Static-SAGE by +7.33 accuracy points; in the second, the effect reverses to -4.05. Both are individually significant under paired analysis, supporting opposite conclusions. The change in the paired effect between executions is highly significant and far larger than within-execution uncertainty. The reversal is consistent across question format, modality, difficulty, and video duration, and both evaluation arms move significantly. Motivated by this failure, we introduce REPAIR (REpeated PAired Inference Reliability), a protocol that repeats the paired comparison and tests whether the method effect changes across executions, separating directional reproducibility from effect-size stability. Applied across accuracy and execution metrics, REPAIR exposes three behaviors-directional reversal, magnitude shift, and effect attenuationand shows that reductions in reasoning turns and visible tool calls do not imply reproducible reductions in primitive computation or latency. The execution-level movement is comparable to, and here larger than, median gain reported by recent agentic VideoQA systems, contextualizing its magnitude without implying those systems are unstable. Significance within a single agentic execution is insufficient evidence that a reported method effect is reproducible.