合成语言模型作为评判语料库中的测试预言机问题:消失、扭曲与验证协议
The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation Protocol
浏览论文内容
中文总结 AI 辅助
研究语言模型作为评判系统中合成语料库的测试预言机问题,指出其负面示例由语言模型生成时存在完整性验证漏洞,通过多语言语料库案例揭示答案截断、偏差扭曲等问题,并提出验证协议供无预言机模式分析师使用。
中文摘要 AI 辅助
对语言模型作为评判系统中的偏差研究通常通过促使语言模型生成一个幻觉答案与事实答案配对,然后将两者呈现给评判者来构建合成语料库。我们报告了一个生成步骤悄然失败的案例,并认为这种失败模式是结构性的而非偶然的。在一个多语言(土耳其语/英语)忠实度评判语料库中,评判和生成调用之间共享的解码预算参数将一个生产者的幻觉答案截断为几个单词。由此产生的项目产生了巨大的、统计上稳健的影响:一位评判者的选择准确率出现了32分的跨语言下降,从N = 50到N = 500都能复制,由一个三层机制解释,并通过受控的生产者交换实验得到证实,但这些都不是真实的。一旦共享参数得到纠正,这种影响就消失到了上限,只有手动阅读原始生成内容,而不是任何聚合统计检查,才能发现故障。第二个测量到的偏差(Markdown格式偏好)没有被伪造,但也因同样的故障而扭曲,其大小以及在一种情况下其符号会随刺激长度而变化,这种模式是聚合指标无法与第一种情况区分开来的。我们使用测试预言机问题来构建潜在的漏洞:其负面示例由语言模型生成的语料库没有机械的方法来验证项目完整性,而通过对黄金答案进行确定性扰动构建的语料库则免费带有项目级预言机。一个积极的对照直接支持了这一说法:注入到基于最小扰动的语料库中的类似故障通过零成本、零人工的黄金到负字符串比较以100%的准确率被捕获。我们最后提出了一个验证协议,该协议源自我们自己的案例,供在我们认为描述了大多数当代多语言语言模型作为评判语料库的无预言机模式下工作的分析师使用。
英文摘要
Studies of bias in LLM-as-judge systems typically build synthetic corpora by prompting an LLM to generate a hallucinated answer to pair with a factual one, then presenting both to a judge. We report a case in which this generation step silently failed, and use it to argue that the failure mode is structural rather than incidental. In a multilingual (Turkish/English) faithfulness-judgment corpus, a decoding-budget parameter shared between judging and generation calls truncated one producer's hallucinated answers to a few words. The resulting items produced a large, statistically robust effect: a 32-point cross-lingual collapse in one judge's selection accuracy, replicated from N=50 to N=500, explained by a three-layer mechanistic account, and confirmed by a controlled producer-swap experiment, none of which was real. The effect vanished to ceiling once the shared parameter was corrected, and only manual reading of the raw generations, not any aggregate statistical check, exposed the fault. A second measured bias (Markdown-formatting preference) was not fabricated but distorted by the same fault, its magnitude and in one case its sign shifting with stimulus length, a mode aggregate metrics cannot distinguish from the first. We frame the underlying vulnerability using the test oracle problem: corpora whose negative examples are LLM-generated carry no mechanical way to verify item integrity, while corpora built by deterministic perturbation of a gold answer carry an item-level oracle for free. A positive control supports this claim directly: an analogous fault injected into a minimal perturbation-based corpus is caught with 100% accuracy by a zero-cost, zero-human gold-to-negative string comparison. We close with a validation protocol, derived from our own case, for analysts working in the oracle-less regime that we argue describes most contemporary multilingual LLM-as-judge corpora.
发表机构
- Mehmet Akif Ersoy University(梅赫梅特·阿基夫·埃尔索伊大学)
机构由 AI 辅助整理,请以论文原文为准。