arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29935stat.MLcs.LG

LLM生成文本在污染下的鲁棒检测

Robust Detection of LLM-Generated Text under Contamination

Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM生成文本检测在编辑和污染下的鲁棒性问题,提出截断似然比检验作为现有统计检测器的简单修改,在多个数据集和基准上显著提升检测鲁棒性。

中文摘要 AI 辅助

我们研究了在编辑和污染条件下对LLM生成文本的检测问题。将人类和机器文本建模为具有Huber污染的有限阶马尔可夫过程,我们在假设下刻画了可靠检测的精确边界。当污染相对于干净源分离足够大时,检测是不可能的。在此边界之下,一组截断似然比检验实现了渐近为零的最坏情况误差。这种构造促使我们将截断作为现有统计检测器的一种简单修改。对于一类广泛的加性分数,我们确定了截断检验一致而原始检验最坏情况功效趋于零的条件。我们在三个数据集、三个生成模型以及RAID基准上评估了七种检测器。截断在两项研究中均提高了鲁棒性,但提升幅度因检测器和污染设置而异。例如,在目标假阳性率为5%时,截断将对数似然-对数秩比(LRR)检测器的真阳性率在受控研究中中位数提高了8.3个百分点,在速率特定和攻击特定的RAID评估中分别提高了2.1和4.3个百分点。

英文摘要

We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test's worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5\%, clipping improves the log-likelihood--log-rank ratio (LRR) detector's true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.

发表机构

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

↑