线性探测提供了对机器生成文本的鲁棒且高效的检测
Linear Probing Provides Robust and Efficient Detection of Machine-Generated Text
查看机构详情
- King’s College London(伦敦国王学院)
- Wikimedia Foundation(维基媒体基金会)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出线性探测器作为鲁棒且样本高效的机器生成文本检测器,在4个基准上优于16种基线,分布外检测AUC提升11,仅需少于100个样本即可接近峰值性能。
中文摘要 AI 辅助
由于机器生成文本(MGT)可能被滥用,区分机器生成文本(MGT)与人类撰写文本(HWT)变得愈发重要。然而,大多数监督式检测器在分布外(OOD)场景下性能常出现退化,且需要大量、多样化的训练集。本研究分析了机器生成文本表示的线性特性与质量,表明简单的线性探测器在性能上优于多种检测器,同时样本效率显著更高。我们首先证明,机器生成文本与人类撰写文本的潜在表示在低维空间中是线性可分的,并通过二者表示质量的系统性差异为这种可分性提供了合理的解释。基于这些发现,我们训练了两种简单线性探测器的变体,在4个基准上对16种基线方法进行评估。这些探测器在分布外检测中持续提升性能(AUC提升11),且仅需少于100个样本即可达到接近峰值的性能。我们表明这种可迁移性源于探测器恢复了跨不同场景通用的共享潜在机器生成文本方向。最后,我们证明探测向量捕捉了“机器生成程度”的连续谱,凸显了其在AI编辑文本细粒度估计方面的潜力。总体而言,本研究揭示了机器生成文本与人类撰写文本在潜在空间中的差异,并证明了线性探测器作为鲁棒且样本高效的机器生成文本检测器的潜力。我们在GitHub上发布了代码。
英文摘要
Distinguishing machine-generated text (MGT) from human-written text (HWT) becomes increasingly important due to potential misuse. However, most supervised detectors often degrade out-of-domain (OOD) and require large, diverse training sets. In this work, we analyze the linearity and quality of MGT representations and show that simple linear probes outperform a wide range of detectors while being substantially more sample-efficient. We first show that MGT and HWT latent representations are linearly separable in low-dimensional space, and provide a plausible explanation for this separability through systematic differences in their representation quality. Motivated by these insights, we train two variants of simple linear probes and evaluate them across 4 benchmarks against 16 baselines. Probes consistently improve OOD detection (+11 AUC), requiring solely ${<}100$ samples to reach near-peak performance. We show that this transferability arises because probes recover a shared latent MGT direction that generalizes across diverse settings. Finally, we demonstrate that probing vectors capture a continuous spectrum of ``machineness'', highlighting their potential for fine-grained estimation of AI-edited text. Overall, our work provides insights into latent-space differences between MGT and HWT and demonstrates the potential of linear probes as as robust and sample-efficient MGT detectors. We release our code on~\href{https://github.com/gerritq/mgt_probes}{github}.