arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01000cs.AI

判断并非枚举:大语言模型生成的可接受集合中的隐性遗漏

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets

Wenhui Chen, Jianlin Chen, Ziyao Lin, Peiji Long, Chi Man Vong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现大语言模型作为评估者时,生成可接受集合的表现远差于判断表现,核心问题是隐性遗漏,通过重写错误预期值可大幅提升修复后的套件产量。

中文摘要 AI 辅助

大语言模型正日益从被评估者转变为评估者:它们编写测试套件、答案密钥、评分标准和奖励函数,以此定义其他系统的正确性。我们评估了这一角色所需的能力,并发现其在该角色通常采用的协议下存在不足——即采用无测试时推理的单样本贪婪生成方式。我们在四种参考构造上开展了实验:两种具备完整有限真值的构造、一种带有加固可执行参考的构造(HumanEval+/MBPP+)、一种带有明确不完整词汇参考的构造(WordNet)。结果显示,模型判断候选是否属于集合的表现远优于其生成集合本身。在针对不完整性的算法构造上,24倍参数范围内的F1分数差距为+0.34至+0.29,且未缩小;在可执行代码上,模型判断的F1分数为0.74至0.90,但生成的套件仅能接纳19%至42%的神谕正确解。一项对照实验定位了缺陷:当要求模型输出谓词而非其外延集时,相同模型的F1分数可达约0.99。该失败并非源于知识缺失或无法指定,而是无法实现规范所诱导的区域。主要错误是遗漏,且这种错误难以审计:过度包含是审核者可质疑的标记,而缺失成员是一种缺失,其发现本身就是生成问题。模型检测植入的过度包含的频率是植入遗漏的6至7倍,且一个包含43227个项目的生产部署中,遗漏优先的失败比例为10:1。将生成的密钥接入RLVR后,相较于精确神谕,其准确率损失1.9个百分点,相较于WordNet基准损失18.5个百分点(六对种子,p=0.031)。对已知正确探针的生成验证器进行门控,可将错误拒绝率从58%至92%降至至多5%,但仅保留5%至39%的套件。通过将每个错误的预期值重写为参考执行返回的值来修复它们,在四个作者族中,产量提高了3.3至10.6倍。

英文摘要

Language models are increasingly promoted from examinees to examiners: they write the test suites, answer keys, rubrics, and reward functions that define correctness for other systems. We measure the capability that role assumes and find it lacking under the protocol the role is usually deployed with, one-shot greedy authoring with no test-time reasoning. Across four reference constructions - two with complete finite truth, one with a hardened executable reference (HumanEval+/MBPP+), one with an explicitly incomplete lexical reference (WordNet) - models judge whether a candidate belongs far better than they author the set itself. On the incompleteness-proof algorithmic construction the gap is +0.34 to +0.29 F1 over a 24x parameter range and does not close; on executable code, models judging at F1 0.74-0.90 author suites admitting only 19-42% of oracle-correct solutions. A control locates the deficit: asked to emit the predicate rather than its extension, the same models reach F1 about 0.99. The failure is not missing knowledge or an inability to specify, but an inability to materialise the region a specification induces. The dominant error is omission, which resists audit: an over-inclusion is a token a reviewer can challenge, a missing member an absence whose discovery is the authoring problem itself. Models detect planted over-inclusions 6-7x more often than planted omissions, and a production deployment of 43,227 items fails omission-first at 10:1. Wired into RLVR, an authored key costs 1.9 points of accuracy against an exact oracle and 18.5 WordNet-relative (six paired seeds, p=0.031). Gating authored verifiers on a known-correct probe cuts false rejection from 58-92% to at most 5%, but keeps only 5-39% of suites. Repairing them instead, by rewriting each wrong expected value to what a reference execution returns, raises yield 3.3-10.6x across four author families.

发表机构

  • University of Macau(澳门大学)
  • South China University of Technology(华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑