评估者是实验的一部分:测量开放式大语言模型(LLM)的一致性
The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity
浏览论文内容
中文总结 AI 辅助
该研究提出实验方案测量开放式LLM一致性,发现错误同行输入会降低修订质量,评估者非中立,锚定校准必要,指出翻转率不足以完整衡量开放式一致性。
中文摘要 AI 辅助
现有关于大语言模型(LLM)一致性的研究大多是在可验证标签下测量离散的答案翻转。开放式修订需要不同的测量策略,因为答案质量是分级的、潜在的,且判断存在不完美性。我们引入了一种实验方案,该方案在合并的主同行条件语料库以及单独构建的分解语料库上实施,使我们能够区分普通重新回答、候选内容暴露、捆绑的同行呈现残差以及评估者对可见同行上下文的方向敏感性。在四个开放权重生成器和三个基准测试中,所有错误的同行输入在每个生成器-数据集单元中都会产生最低质量的修订。对相同答案的盲评和知情评分也因评估者而异:一名评估者转向同行认可的立场,两名评估者远离该立场,一名评估者大致中立,GPT-4o和GPT-5.4-mini的评估同样非中立。最后,一项锚定评估显示,简洁的正确锚点可能被误读,频率足以破坏潜在量表,除非明确检查校准。这些结果支持四个结论:翻转率不足以作为开放式一致性的完整测量,错误的同行会损害开放式修订,评估者并非中立,锚定校准是必要的。
英文摘要
Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly. We introduce an experimental protocol implemented across a pooled main peer-condition corpus and separately constructed decomposition corpora, allowing us to separate ordinary re-answering, candidate-content exposure, a bundled peer-presentation residual, and directional judge sensitivity to visible peer context. Across four open-weight generators and three benchmarks, all-wrong peer input produces the lowest-quality revisions in every generator-dataset cell. Blind and informed ratings of identical answers also differ by evaluator: one judge shifts toward the peer-endorsed position, two shift away, one is approximately neutral, and GPT-4o and GPT-5.4-mini audits are likewise non-neutral. Finally, an anchor audit shows that terse correct anchors can be misread often enough to destabilize the latent scale unless calibration is checked explicitly. These results support four conclusions: flip rates are insufficient as a complete measure of open-ended conformity, wrong peers harm open-ended revision, evaluators are not neutral, and anchor calibration is necessary.
发表机构
- Illinois Institute of Technology(伊利诺伊理工大学)
机构由 AI 辅助整理,请以论文原文为准。