反事实敏感性不等于可修复性:针对视频证据的重播探针审计
Counterfactual Sensitivity Is Not Repairability: Auditing Replay Probes for Video Evidence
浏览论文内容
中文总结 AI 辅助
本研究提出黑盒反事实探针 CARVE,通过对比 SHAM 与 DESTROY 重播下的答案变化,审计视频智能体的证据依赖,在 LVBench 上取得准确率提升,验证了反事实敏感性与可修复性的差异。
中文摘要 AI 辅助
使用工具的视频智能体在回答前会检索视觉证据,但最终答案并不强制依赖所检索的内容。自然的黑盒测试是反事实的:破坏智能体检索到的帧的语义内容,并将答案与在这些相同帧上重新执行相同流程的匹配 sham 进行比较。我们引入 CARVE,一种黑盒反事实探针,用于比较匹配的 SHAM 和 DESTROY 重播下的答案变化。在冻结的 VideoExplorer 式智能体上进行的三次独立 k=3 运行中,DESTROY 改变答案的频率比 SHAM 高 29.3 个百分点,产生了显著且可重复的总体效应。问题级分数稳定性较低,将重播预算从 k=3 增加到 k=10 会减少平局,但削弱了原始的零阈值路由策略。在 k=3 时,CARVE 从 1258 个 LVBench 问题中选择了 538 个,准确率提高了 3.26 个百分点,且比大多数匹配的随机子集具有更高的后备产量。该分数与标注的时间覆盖率仅显示弱关联,因此 CARVE 最好被理解为一种路由信号,而非直接的 grounding 分类器。我们的实现可在该 https URL 获取。
英文摘要
Tool-using video agents retrieve visual evidence before answering, but the final answer is not forced to depend on what was retrieved. The natural black box test is counterfactual: destroy the semantic content of the frames the agent retrieved and check whether the answer changes, against a matched sham that re-executes the identical pipeline on those same frames. We introduce CARVE, a black-box counterfactual probe that compares answer changes under matched SHAM and DESTROY replays. Across three independent k=3 runs on a frozen VideoExplorer-style agent, DESTROY changes the answer 29.3 percentage points more often than SHAM, yielding a large and reproducible aggregate effect. Question-level scores are less stable, and increasing the replay budget from k=3 to k=10 reduces ties but weakens the original zero-threshold routing policy. At k=3, CARVE selects 538 of 1,258 LVBench questions and improves accuracy by 3.26 points, with higher fallback yield than most matched random subsets. The score shows only a weak association with annotated temporal coverage, so CARVE is best understood as a routing signal rather than a direct grounding classifier. Our implementation is available at https://github.com/KurbanIntelligenceLab/CARVE.