AI 中文总结
该研究构建了RENDEQ生成科学图形渲染等价集,测量VLMs的共识-准确性耦合,发现重渲染优于重采样,共识在多数模型上优于基线,微调共识会降低准确性,共识仅在错误扩散阈值以上才证明正确性。
AI 中文摘要
模型在受扰动输入上的共识既被用作无标签可靠性信号,也被用作自训练目标,前提是共识能追踪正确性。这种耦合关系很少被直接测量:自然图像的扰动仅在假设下保留语义,且没有精确答案键来定位错误。科学图形消除了这两个障碍:图形由程序从数据中绘制,因此重绘它会产生在构造上语义等价且共享程序精确答案的图像。我们构建了RENDEQ,一种此类渲染等价集的生成器,并在三个开放权重VLMs上测量该耦合关系,在三个独立实例中检查每个发现。重渲染在准确性和可靠性上均优于重采样。在三个模型中的两个上,共识优于携带证据的基线(平均标记对数概率),在第三个模型上持平,这逆转了归因于渲染管道故障的中间错误复现。这种背后的分散性集中在一个风格因素(绘图库)上,其影响是下一个最大因素的两倍多,且比噪声基底高一个数量级。对模型自身跨渲染共识进行微调会产生相反结果:在五次复现运行中,准确性均下降,这与已发表的自然图像结果符号相反。共识仅在超过模型错误扩散程度设定的阈值时才能证明正确性,而奖励共识的目标恰好破坏了这种扩散性。
英文摘要
A model's agreement across perturbed inputs is used both as a label-free reliability signal and as a self-training target, on the premise that agreement tracks correctness. That coupling is rarely measured directly: natural-image perturbations preserve meaning only by assumption, and no exact answer key localizes errors. Scientific figures remove both obstacles, a figure is drawn from data by a program, so redrawing it yields images that are semantically equivalent by construction and share a programmatically exact answer. We build RENDEQ, a generator of such render-equivalence sets, and measure the coupling on three open-weight VLMs, checking every finding across three independent instantiations. Re-rendering beats resampling on both accuracy and reliability. Agreement beats an evidence-carrying baseline, mean token log-probability, on two of three models and ties on the third, reversing an intermediate, buggy replication traced to a rendering-pipeline failure. The dispersion behind this is concentrated in one style factor, the plotting library, more than double the next-largest factor and an order of magnitude above the noise floor. Fine-tuning on the model's own cross-render consensus inverts: accuracy falls in every one of five replication runs, the opposite sign to published results on natural images. Agreement certifies correctness only above a threshold set by how diffuse a model's errors are, and an objective that rewards agreement destroys exactly that diffuseness.