arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

审稿人是否仍奖励词汇复杂度?对124K篇ICLR评审中偏好漂移的冻结评审员研究

Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews

Jiabin Zheng

arXiv 2609.08475首次发表:更新:

发表机构

School of Computer Science, Peking University(北京大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过冻结评审员分离投稿构成与评审者偏好变化,发现人工评审对词汇复杂度的奖励从+0.142降至-0.015,而LLM评判者仍按早期速率奖励,揭示偏好漂移及校准LLM的失准问题。

AI 中文摘要

大型语言模型已大幅降低了生成词汇繁复文本的成本,而同行评审者是否仍对此给予奖励,这是一个关于评审者而非文本的问题。当写作线索与评审分数之间的关联在不同年份间发生变化时,可能是评审者发生了变化,可能是投稿发生了变化,或者两者兼有,而仅凭分数对文本的回归无法区分这些情况。我们通过冻结评审员将两者分离:对2018年至2025年ICLR投稿的81,850条机器评审,全部在2025年2月至4月这一窗口期内,使用同一模型家族和同一提示词生成,因此其逐年系数仅反映投稿构成的变化,而人工减去冻结的差异趋势则识别出评审者的偏好漂移。在32,638篇投稿的124,615条人工评审中,非领域词汇复杂度的人工系数从+0.142降至-0.015,而冻结评审员从+0.080升至+0.082;三重差分估计为-0.0100(q=0.013),且通过相同规范进行的四十个随机词表安慰剂检验结果集中于零。人工评审者仍奖励句子长度变异性,而冻结评审员从未对此有所体现,同时冻结评审员仍以早期速率对词汇复杂度给予回报。每项主张均通过错误发现率控制和区间排除的双重门槛,并报告了未通过对抗性再检验的发现。评审者贬低了生产成本骤降的线索,正如可操纵信号模型所预测的那样;校准至历史人工偏好的LLM评判者继承了早期的时间表并逐渐偏离对齐,而其与人工在总分上的一致性仍保持一般水平。

英文摘要

Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may have changed, or both, and a regression of scores on text cannot say which. We separate the two with a frozen rater: 81,850 machine reviews of ICLR submissions from 2018 to 2025, all generated in one February-April 2025 window with one model family and one prompt, so that its year-to-year coefficients track submission composition alone and the human-minus-frozen trend difference identifies reviewer preference drift. On 32,638 submissions with 124,615 human reviews, the human coefficient on non-domain lexical complexity falls from +0.142 to -0.015 while the frozen rater moves from +0.080 to +0.082; the three-way difference-in-differences is -0.0100 (q=0.013), and forty random-wordlist placebos through the same specification centre on zero. Humans still reward sentence-length variability, which the frozen rater never registers, while the frozen rater still pays for lexical complexity at its earlier rate. Every claim is held to a double gate of false-discovery control and interval exclusion, and the findings that failed adversarial re-testing are reported. Reviewers discounted a cue whose production cost collapsed, as models of manipulable signals prescribe; an LLM judge calibrated to historical human preferences inherits the earlier schedule and drifts out of alignment while its agreement with humans on totals stays ordinary.

Comments23 pages, 8 figures, 11 tables. Code and the machine-readable records behind every number: https://github.com/Biajin-PKU/frozen-rater-drift

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑