arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我们是否在测量预期?程序性视频评估中特权信息的审计

Are We Measuring Anticipation? Auditing Privileged Information in Procedural Video Evaluation

Mahsa Mohammadi, Sareh Rowlands

arXiv 2610.03826首次发表:更新:

发表机构

University of Exeter(埃克塞特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究审计程序性视频评估协议,发现即使无时间泄漏也可能存在特权信息差距,导致能力声明与实测构念不符,并提出可复用的四条件审计方法。

AI 中文摘要

基准分数为关于被评估能力的声明提供了依据。我们审计的是评估协议所支持的推断,而不仅仅是预测模型本身。以程序性动作预期作为受控案例研究,我们研究了一个更广泛的评估有效性失败模式:一个协议可以保持时间因果性且没有经典的目标泄漏,同时仍然提供一种特权的中间表示。这种失败不仅仅是乐观的准确性,而是所声称的能力与实际测量的构念之间的不匹配。在Breakfast数据集上,匹配的识别器历史和GT历史方案得分分别为30.3%和62.3%(Delta_PH(R_BA) = +32.0,95% Student-t置信区间[+24.1, +39.8])。一个固定检查点的2×2干预隔离出+14.6分的测试时神谕对比;一个匹配的无视频来源探针产生+26.8分的差距;一个与边界无关的查询/历史控制保留了+19.4分的差距。我们将这种差异称为特权信息差距,其定义相对于指定的非神谕恢复管线,并提出一个可复用的四条件审计。在五个评估设置中,该差距是异质的;一个标准的50 Salads长期预期重新实现为审计特定的下一动作构造之外提供了谨慎的支持。

英文摘要

Benchmark scores license claims about the capabilities being evaluated. We audit the inference licensed by an evaluation protocol, rather than the predictive model alone. Using procedural action anticipation as a controlled case study, we study a broader evaluation-validity failure mode: a protocol can remain temporally causal and free of classical target leakage while still supplying a privileged intermediate representation. The failure is not merely optimistic accuracy, but a mismatch between the capability claimed and the construct actually measured. On Breakfast, matched recognizer-history and GT-history regimes score 30.3% and 62.3% (Delta_PH(R_BA) = +32.0, 95% Student-t interval [+24.1, +39.8]). A fixed-checkpoint 2 x 2 intervention isolates a +14.6-point test-time oracle contrast; a matched no-video provenance probe yields a +26.8-point gap; and a boundary-independent query/history control retains a +19.4-point gap. We refer to this discrepancy as a privileged-information gap, defined relative to a specified non-oracle recovery pipeline, and propose a reusable four-condition audit. Across five evaluation settings the gap is heterogeneous; a standard 50 Salads long-term-anticipation re-implementation provides cautious support beyond the audit-specific next-action construction.

Comments17 pages, 6 figures. NeurIPS 2026 Workshop on TAE (Trust-AI-Eval): Can We Trust AI Evaluation?

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑