一致性存在可计算的盲区:视觉-语言图表阅读的无标签可靠性交换理论
Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading
AI总结:
该研究提出视觉-语言图表阅读的无标签可靠性交换理论,构建无标签、无需训练的检测器,发布REND-EQUIV数据集,揭示一致性盲区可计算,循环重标记可提升性能,解释排序反转现象。
AI中文摘要:
视觉-语言模型的无标签可靠性基于不变性:对输入进行扰动时,可靠读取器的答案不应改变。这存在一个已知盲区:系统误读在扰动后仍存在,且可被判定为错误,我们证明该盲区是可计算的,而非仅为真实存在:当两个操作可交换时,错误对编辑不可见,因此一组编辑无法触及的错误构成其联合中心化子,该集合会随编辑增加而缩小,且可被明确写出而非猜测。我们利用互补的等变性:编辑图表数据时,正确答案必须按可计算的量变化。两个匹配编辑对仿射读取错误是完备的;对于标签置换,任何交换编辑集合都不完备,而循环重标记可填补大部分该缺口。我们将该理论实例化为等变性-一致性得分(Equivariance-Consistency Score),这是一种无标签、无需训练的检测器,并发布了REND-EQUIV数据集,其在相同数据上配对了匹配的不变性和等变性集合。预测的排序在三个模型和一个手动标记的总体上成立,该总体不受其选择方式中的一个循环性影响;第二种不变性族方法证实,该盲区属于关系而非任何实现;循环重标记在匹配的真实样本上实现了其预测的增益。相同的特征解释了分类器变态测试文献中报道的这种排序反转:可检测性是关系和错误类别的联合属性,绝不仅是关系本身的属性。
英文摘要:
Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at. We act on the complementary relation, equivariance: edit a figure's data and the correct answer must change by a computable amount. Two matched edits are provably complete for affine reading errors; no suite of swap edits is complete for label permutations, and cyclic relabeling closes most of that gap. We instantiate the theory as the Equivariance-Consistency Score, a label-free, training-free detector, and release REND-EQUIV, pairing matched invariance and equivariance sets over identical data. The predicted ordering holds across three models and a hand-labeled population immune to the one circularity in how it is selected; a second invariance-family method confirms the blind spot belongs to the relation, not to any implementation; and cyclic relabeling delivers its predicted gain on a matched real sample. The same characterization explains a reported inversion of this ordering in the classifier metamorphic-testing literature: detectability is a joint property of the relation and the fault class, never of the relation alone.