发表机构
Seoul National University; University of Minnesota(首尔大学; 明尼苏达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成论文中的科学垃圾问题,提出SciSlopBench基准和SciSlopHarness框架,通过六项指标检测并以证据为基础修订,显著缩小AI与人类论文差距。
AI 中文摘要
AI生成的内容,通常被称为AI垃圾,在各地日益普遍,尤其是在学术界。然而,AI生成的科学论文中的垃圾具有更复杂的模式,现有的基于标记的AI检测器难以轻易识别。这类论文的每个部分看起来都合理,但连接各部分之间的科学推理却断裂了,这可能会误导读者对工作的评估。我们通过结构、论证和伪影六个指标将这些失败作为科学垃圾进行基准测试。我们构建了SciSlopBench,包含390篇AI生成的论文,主要来自计算机科学,但也涵盖生命科学、社会科学和自然科学,每篇论文都与一篇在研究问题和贡献类型上匹配的人类撰写论文配对。我们的指标在每对中识别AI论文的准确率为85.9%,而Binoculars的准确率为68.7%。更高的科学垃圾水平伴随着更低的ICLR评分,并且在2017年至2025年的每一年中,都能以高于偶然水平的概率区分被拒绝和被接受的论文。然而,减少这些模式并非像直接优化指标那样简单。因此,我们提出了SciSlopHarness,一个框架级框架,指导固定的LLM仅在实验记录支持更改的地方修改垃圾。虽然标准修订会留下残余垃圾,而直接感知垃圾的提示会引发奖励黑客行为,但SciSlopHarness在最强修订基线上将剩余的AI-人类差距减少了63%,且无需人类参考目标。总体而言,我们证明了AI生成的科学论文在其全局推理中留下了基本痕迹,并且负责任的缓解需要严格的证据基础,而不仅仅是散文润色。
英文摘要
AI-generated content, often called AI slop, is increasingly common everywhere, particularly in academia. Slop in AI-generated scientific papers, however, has more complex patterns that cannot be easily detected by existing token-based AI detectors. Each part of such a paper looks plausible while the scientific reasoning that connects the parts breaks down, which can mislead how readers assess the work. We benchmark these failures as scientific slop through six measures across Structure, Argument, and Artifacts. We construct SciSlopBench with 390 AI-generated papers, mostly in computer science but spanning the life, social, and natural sciences, each paired with a human-written paper matched by research problem and contribution type. Our measures identify the AI paper in each pair with 85.9% accuracy, compared with 68.7% for Binoculars. Higher scientific slop accompanies lower ICLR ratings and distinguishes rejected from accepted papers above chance in every year from 2017 to 2025. Reducing these patterns, however, is not as simple as directly optimizing the measures. We therefore propose SciSlopHarness, a harness-level framework that guides a fixed LLM to revise slop only where the experiment records support the change. While standard revisions leave residual slop and direct slop-aware prompting triggers reward hacking, SciSlopHarness reduces the remaining AI-human gap by 63% over the strongest revision baseline without requiring human reference targets. Overall, we demonstrate that AI-generated scientific papers leave fundamental traces in their global reasoning, and that responsible mitigation demands strict evidentiary grounding rather than mere prose refinement.
Comments28 pages, 6 figures, 13 tables. Project page: https://yerimoh.github.io/scientific-slop-demo/