arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少即是多:理解AI生成文本检测中词元过滤的有效与失效情形

When Less is More: Understanding When Token Filtering Helps and Fails in AI-generated Text Detection

Xiaoyang Han, Lvxiaowei Xu, Ming Cai

arXiv 2608.29903首次发表:更新:

发表机构

College of Computer Science and Technology; College of Artificial Intelligence; Zhejiang University(计算机科学与技术学院; 人工智能学院; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对AI生成文本检测中词元过滤的共识,通过实证与理论分析,发现仅保留40%词元可获最优性能,且过滤效果因源LLM强弱而异,揭示了词元级检测的双向权衡。

AI 中文摘要

大型语言模型(LLM)的快速发展使得AI生成文本的检测愈发关键。现有的零样本检测器假设更多词元级证据会带来更可靠的检测结果,然而我们的实证研究对这一共识提出了挑战:更少的词元有时效果更好,仅保留40%即可达到最优性能,但这种益处并非普遍存在。我们利用熵差距分数(EGS)引入前k个累积概率过滤作为诊断工具,在三种代表性场景中,过滤表现出截然不同的行为。我们通过典型集理论分析EGS,并通过熵校准和分布分析量化其动态变化,发现过滤对弱源LLM有帮助,其中低熵词元有害;但对强源LLM则失效,低熵词元并无明显危害。本研究首次系统分析表明,部分词元不仅无信息,还因熵校准偏差具有系统性危害,揭示了词元级检测中的双向权衡。

英文摘要

The rapid advancement of large language models (LLMs) has made AI-generated text detection increasingly critical. Existing zero-shot detectors assume that more token-level evidence leads to more reliable detection. However, our empirical study challenges this consensus: fewer tokens sometimes work better, retaining only 40% can yield optimal performance, yet this benefit is not universal. Using the Entropy Gap Score (EGS), we introduce top-$k$ cumulative probability filtering as a diagnostic probe. Across three representative settings, filtering exhibits strikingly different behaviors. We analyze EGS via typical set theory and quantify its dynamics through entropy calibration and distribution analysis. We find that filtering helps for weak source LMs, where low-entropy tokens are harmful, but fails for strong source LMs, where they are not notably harmful. Our work provides the first systematic analysis showing that some tokens are not merely uninformative but systematically harmful due to entropy miscalibration, revealing a two-sided trade-off in token-level detection.

CommentsAccepted to EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑