arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08975cs.CLcs.AI

如何对AI审稿人进行修辞奖励黑客攻击?剖析基于AI的同行评审中的修辞敏感性

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对基于AI的同行评审,构建受控语料库测试修辞选择对AI审稿判断的影响,明确修辞敏感性的层级结构及工作流程效果,为开发鲁棒评估系统提供依据。

中文摘要 AI 辅助

随着大型语言模型越来越多地参与科学评估,我们研究一种潜在的奖励黑客攻击形式:在报告的科学内容保持不变的情况下,修辞选择如何影响AI审稿人的判断,以及这些影响在不同评估条件下如何变化。我们构建了一个受控语料库,包含4200篇完整论文手稿,这些手稿源自120篇匿名的ICLR 2026投稿。两名大型语言模型(LLM)重写器从相反方向改变六个修辞维度,五名LLM审稿人在标准和严格协议下评估得到的手稿。我们还测试了联合、递归和审稿人引导的重写。我们的结果表明,修辞敏感性是结构化的,而非均匀的。证据框架和新颖性立场在整体评估中产生最大的正负对比,范围框架形成较弱的第二梯队;其余维度的影响较小或较不稳定。这种层级在人类评估的质量水平中持续存在,但分数变动强烈依赖于AI审稿人的原始分数:较低分数倾向于上升,较高分数倾向于下降,方向对比在中间范围最清晰。更复杂的工作流程无法可靠地产生更大收益。联合重写强烈依赖于重写器,审稿人引导并未始终优于无引导的二次重写,重复重写产生递减的、依赖于配置的回报。在所有条件下,重写器主要决定相反变体之间的差异,而审稿人决定其分数影响的幅度和符号。严格评审使平均OA降低1.36分,且未始终改变修辞敏感性。这些发现确定了修辞呈现何时影响AI科学评审,并催生了对科学写作中内容保留变化具有鲁棒性的评估系统。

英文摘要

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.

发表机构

  • University of Maryland(马里兰大学)
  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

↑