在人类与大语言模型合作撰写的文本中检测大语言模型生成的令牌
Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
浏览论文内容
中文总结 AI 辅助
针对人类与大语言模型合作撰写文本中难以检测大语言模型生成内容的问题,提出在令牌级别运行、基于现有分数并平滑分数降低变异性的新方法,在合成与真实数据集上性能强,还部署了网站。
中文摘要 AI 辅助
人类与人工智能协作写作的兴起,使得对支持在混合撰写文档中定位可能由大语言模型生成内容的细粒度检测方法的需求日益增长。现有检测大语言模型生成文本的方法主要集中在文档级分类,无法识别文本中哪些部分是由大语言模型生成的。本文介绍了一种新方法来满足这一迫切需求。我们的方法在令牌级别(现代语言模型的自然单元)运行,并基于现有的令牌级检测分数。关键思想是平滑相邻令牌分数以降低其变异性,同时使用自适应Lepski型规则根据局部作者结构选择带宽。我们的方法易于实现,不需要令牌级标记数据进行训练。理论上,我们刻画了这种权衡,并表明所提出的方法在估计潜在信号方面实现了良好的均方误差性能。实证上,我们在合成数据集和真实数据集上针对广泛的基线展示了我们方法的强大性能。我们还部署了一个公开访问的网站来实现这些方法。
英文摘要
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs. This paper introduces a new method to address this urgent need. Our method operates at the token level, the natural unit of modern language models, and builds on existing token-level detection scores. The key idea is to smooth adjacent token scores to reduce their variability, while using an adaptive Lepski-type rule to select the bandwidth according to the local authorship structure. Our method is simple to implement and does not require token-level labeled data for training. Theoretically, we characterize this trade-off and show that the proposed method achieves favorable mean square error performance in estimating the underlying signal. Empirically, we demonstrate strong performance of our method against a wide range of baselines in both synthetic datasets and a realistic dataset. We deploy a publicly accessible website that implements the methods as well.
发表机构
- School of Mathematics, University of Birmingham(伯明翰大学数学学院)
- School of Statistics and Data Science, Shanghai University of Finance and Economics(上海财经大学统计与数据科学学院)
- Department of Statistics, The London School of Economics and Political Science(伦敦政治经济学院统计系)
机构由 AI 辅助整理,请以论文原文为准。