发表机构
Tokyo Metropolitan University; National Institute of Informatics(东京都立大学; 信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究量化了LLM版本更迭对科研写作筛查的影响,发现仅基于旧版本训练的检测器会在模型代际边界处失效,建议科研诚信政策需随LLM版本更新重新验证检测器准确性。
AI 中文摘要
期刊和会议已开始筛查提交的手稿中是否有使用大语言模型(LLM)撰写的文本。这种筛查的可靠性依赖于针对固定一组LLM版本的基准评估,而实际使用的版本却在不断变化。本文量化了这种LLM版本更迭对科研手稿筛查的影响。我们将来自《美国国家科学院院刊》的4000篇ChatGPT出现前的摘要,与2023年6月至2026年8月期间三家厂商发布的23个LLM版本对这些摘要的重写版本进行配对。随后,我们在多种维护场景下训练检测器,场景范围从每出现一个新版本就重新训练检测器,到仅训练一次且永不更新。仅基于某厂商过去版本训练的检测器会在模型代际边界处失效:在校准为错误标记1%的人类撰写摘要的情况下,它们能捕获边界前最尖锐处的99%以上重写版本,而边界后仅能捕获3.8%的重写版本。基于较新版本训练的检测器也会遗漏早期版本的重写版本。版本间的词汇差异在很大程度上决定了检测的可迁移与失效位置。在我们模拟的两种筛查场景中,覆盖全部23个版本的筛查要么会标记八分之一的人类撰写摘要,要么会遗漏三分之一的最新版本重写版本。实际上,某商业检测器遗漏了边界后最尖锐处之后版本的大部分重写版本,同时几乎未标记任何人类撰写摘要。因此,科研诚信政策应将检测器的基准准确性视为临时的,需在每次LLM发布(包括早期版本)时重新验证。
英文摘要
Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs). The reliability of this screening rests on benchmark evaluations against a fixed set of LLM versions, while the versions in actual use keep changing. Here we quantify how this LLM turnover affects the screening of scientific manuscripts. We paired 4,000 pre-ChatGPT abstracts from the Proceedings of the National Academy of Sciences with their rewrites by 23 LLM versions from three vendors, released between June 2023 and August 2026. We then trained detectors under maintenance scenarios ranging from a detector retrained on every new version to one trained once and never updated. Detectors trained only on a vendor's past versions can collapse at the boundaries between model generations: calibrated to falsely flag 1% of human-written abstracts, they catch above 99% of rewrites just before the sharpest boundary and 3.8% just after it. Detectors trained on later versions can also miss rewrites of earlier ones. Vocabulary differences between versions largely track where detection transfers and where it fails. In the two screening scenarios we simulated, screens covering all 23 versions either flagged one in eight human-written abstracts or missed one in three rewrites of the newest version. Indeed, a commercial detector missed most rewrites of the version just after the sharpest boundary while flagging almost no human-written abstracts. Research-integrity policy should therefore treat the benchmark accuracy of a detector as provisional, to be re-verified with every LLM release, including earlier versions.