使用LLM检测LLM生成文本:跨代分析
Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis
- University of Kent(肯特大学)
- Shanghai Jiao Tong University(上海交通大学)
- University of Birmingham(伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究系统评估15个跨代LLM作为生成器和检测器的表现,发现检测效能主要取决于检测器能力而非生成器来源,且自检测无系统性优势,同时揭示了代际偏差转移现象。
中文摘要 AI 辅助
自动检测LLM生成的文本(LGTs)至关重要,然而专门的检测器往往难以在不同领域和模型间泛化。虽然通用LLM提供了灵活的零样本作者身份分类并附有解释性理由,但其检测行为,特别是跨模型代际的自检测与交叉检测,仍鲜为人知。我们系统评估了15个覆盖三代模型的LLM,既作为生成器也作为检测器。使用包含1,000篇人类撰写文本和15,000篇LGTs(每模型1,000篇)的基准,我们收集了超过233,000个二元分类结果及自然语言解释。结果表明,检测效能主要由检测器能力而非生成器来源驱动,尽管较新生成器的输出仍明显更难检测。关键的是,统计比较显示,各模型在自检测方面并无系统性优势或劣势。错误分析进一步揭示了代际偏差转移:第一代检测器对LGTs检测不足(高假阴性率),第二代检测器过度标记人类文本(高假阳性率),而最新模型实现了平衡的权衡。最后,我们强调了不同LLM在应用文本线索以证明其决策时存在显著不一致性。代码:此https URL。
英文摘要
Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self-detection versus cross-detection across model generations, remains poorly understood. We systematically evaluate 15 LLMs spanning three model generations as both generators and detectors. Using a benchmark of 1,000 human-written texts and 15,000 LGTs (1,000 per model), we collected over 233,000 binary classifications alongside natural-language explanations. Our results reveal that detection efficacy is primarily driven by detector capability rather than generator provenance, although outputs from newer generators remain notably harder to detect. Crucially, statistical comparisons show no systematic advantage or disadvantage for self-detection across models. Error analysis further exposes generational bias shifts: first-generation detectors under-detect LGTs (high false-negative rates), second-generation detectors over-flag human texts (high false-positive rates), and the latest models achieve balanced trade-offs. Finally, we highlight significant inconsistencies in how different LLMs apply textual cues to justify their decisions. Code: https://github.com/hyyuan/detect-llm-generated-texts.