发表机构
LEX AI(LEX AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究揭示法律多项选择基准存在仅依赖选项的可解性漏洞,模型可通过答案位置偏好盲选得分,且这种漏洞仅属于特定模型,不同数据集格式和模型能力会影响漏洞表现,研究发布了相关语料库与工具。
AI 中文摘要
多项选择基准的评分依据是模型是否选择了正确选项,而非是否真正理解了问题本身。衡量这一差距需要谨慎:若模型对大多数题目都选A,那么无论正确选项是A还是其他,其得分都会高于随机水平;当正确选项并非A时,这种表现会被误认为是对问题的理解。我们在UA-JudgeExam数据集上进行了衡量,该数据集包含乌克兰高等法官资格委员会发布的11990道四选一题目及官方标准答案。仅向模型展示选项而不提供问题时,Claude Haiku 4.5的得分为0.383,高于随机水平,且这种漏洞具有集中性:在所有8种选项排列顺序下,有11.8%的题目被模型盲选正确,而随机水平仅为0.2道题。这并非是引用匹配:在搜索280059版乌克兰立法后,仅能匹配到0.128道题。排除这些题目后,剩余8128道题中,用于门控的模型自身得分为0.204,未参与筛选的GPT-5.6在隐藏问题的情况下仍能答对其中51.5%的题目。对整个数据集上的12个保留模型进行评分并减去各自的答案位置偏好后,仅有两个模型保留了超额得分:GPT-5.6为+0.265,Sonnet 4.6为+0.081。若不扣除这种偏好,排名会产生误导:Llama 3.1 8B盲选得分为0.292,仅低于上述两个模型,这完全是因为它对92%的题目都选A。门控机制确实筛选出了真实存在的现象:在被拒绝的题目上,12个模型中有11个的得分在0.518至0.789之间,每个区间都明显高于同一模型在保留题目上的得分。但这种信号仅属于门控所用的模型,基于此进行筛选无法向上迁移。在400题的样本中,这种信号不可见,9个模型的表现均处于“统计随机水平”。重写干扰项的尝试则过度修正至0.168,低于随机水平,仍具有可利用性。对LEXam数据集进行相同探测时,结果为随机水平:该数据集中每个选项都嵌入题干,且没有选项长度超过33个字符。题目格式决定了问题是否会出现,模型能力决定了漏洞被利用的程度。我们发布了该语料库、模型预测结果及测试工具。
英文摘要
Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at A, and reads as recognition when it is not. We measure it on UA-JudgeExam: 11,990 four-option items with official keys, published by Ukraine's Higher Qualification Commission of Judges. Shown the options and no question, Claude Haiku 4.5 scores 0.383 against chance, and the leak is concentrated: 11.8% of items are answered blind on all eight option orders, against 0.2 items expected by chance. It is not quotation: search over 280,059 editions of Ukrainian legislation recovers 0.128. Gating those out retains 8,128 items, on which the gating model itself now scores 0.204, and GPT-5.6, which took no part in the selection, still answers 0.515 of them with the question hidden. Scoring twelve held-out models on the whole set and subtracting each one's answer-position habit, only two keep an excess: GPT-5.6 at +0.265, Sonnet 4.6 at +0.081. Without it the ranking misleads: Llama 3.1 8B scores 0.292 blind, above every model but those two, purely by answering A to 92% of items. The gate does select something real: on the items it rejected, eleven of twelve models score 0.518-0.789, every interval clear of what the same model scores on the items it kept. But that signal is one model's, and filtering on it does not transfer upward. Neither is visible on a 400-item sample, where nine models read as "statistically at chance". Rewriting distractors instead overshoots to 0.168, below chance and as exploitable. The same probe on LEXam returns chance: every option there points into the stem, none longer than 33 characters. Item format decides whether the problem can arise; capability decides how much is extracted. We release the corpus, the predictions and the harness.
Comments21 pages, 4 figures. Dataset, model predictions and code at https://huggingface.co/datasets/overthelex/ua-judge-exam