arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06190cs.CV

一次扰动是不够的:行为式AI评估的可辨识性与盲基线

One Perturbation Is Not Enough: Identifiability and Blind Baselines for Behavioral AI Evaluation

Rasul Khanbayov, Mariam Sohail, Ahmed Abdala, Hasan Kurban

首次发表
浏览论文内容

中文总结 AI 辅助

本研究证明行为式AI评估至少需要n次扰动才能辨识n个输入的使用,并推导出盲基线分数;实验表明三个视觉-语言模型均低于盲界,无法与忽略视频的模型区分。

中文摘要 AI 辅助

行为式评估通过扰动输入并读取输出中由此引起的变化,来验证系统是否使用了该输入。我们证明,此类认证所需的扰动数量是固定的,且报告单次扰动无法满足该要求。当响应比率是策略而非测试项的性质时,行为记录是对指数向量的线性测量,该向量记录输出对每个输入的依赖程度,因此当扰动的对数张成输入空间时,扰动恰好能辨识输入的使用情况。对于n个输入,至少需要n次扰动;不完整的设计恰好混淆了沿其设计矩阵核方向不同的策略,且锐化扰动永远不能替代添加独立扰动。我们还以闭式推导出此类测试对完全不读取输入的策略所给出的分数,该分数远非零,且我们调查的探针均未报告此分数。在正确答案由量纲分析确定的场景中,我们针对三个从视频中报告物理量的视觉-语言模型,运行了一组完整的三个扰动。所有三个模型的得分都远低于其自身的盲界,而非高于盲界,因为每个模型都默认采用一组较小的取整校准值之一,这些值从未与量表所声称的数值匹配;没有一个模型将其重标号响应移动哪怕一个指数,且没有一个模型与指示忽略视频的同一模型可区分。

英文摘要

Behavioral evaluations perturb an input and read the induced change in the output in order to certify that a system uses that input. We show that the number of perturbations such a certificate requires is fixed, and that reporting a single perturbation cannot supply it. Where a response ratio is a property of the policy rather than of the test items, the behavioral record is a linear measurement of an exponent vector recording how much the output depends on each input, so perturbations identify input use exactly when their logarithms span the input space. At least $n$ are needed for $n$ inputs, an incomplete design confuses precisely the policies differing along the kernel of its design matrix, and sharpening a perturbation never substitutes for adding an independent one. We also derive in closed form the score such a test awards a policy that reads nothing, which is far from zero and which none of the probes we survey reports. Instantiating this where the correct response is fixed by dimensional analysis, we run a complete identifying set of three perturbations on three vision--language models reporting a physical quantity from video. All three score far below their own blind bound rather than above it, because each defaults to one of a small set of round calibration values that never matches what the scale asserts; none moves its relabeling response by a single exponent, and none is separable from the same model instructed to ignore the video.

↑