EASE:用于规避AI生成文本检测器的熵自适应分布塑形
EASE: Entropy-Adaptive Distribution Shaping for Evading AI-generated Text Detectors
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出EASE框架,利用源LLM的预测熵自适应调整解码分布,无需训练或检测器反馈即可有效规避AI文本检测器,且保持文本质量与推理效率。
AI中文摘要:
AI生成文本(AIGT)检测可能对源大语言模型(LLM)的解码选择敏感。我们观察到,扰动下一令牌的对数几率或调整采样温度可以降低检测性能,这为检测器对解码时分布变化的脆弱性提供了明确信号。基于这一观察,我们提出了EASE(用于规避的熵自适应分布塑形),这是一种无需训练且与检测器无关的框架,用于规避AIGT检测器。EASE直接从源LLM的下一令牌分布计算预测熵,并利用它自适应调整对数几率扰动和采样温度,无需检测器反馈或模型微调。在三个源LLM和多个检测器上的实验表明,检测性能持续降低,同时文本质量下降可忽略不计,推理开销也可忽略不计。
英文摘要:
AI-generated text (AIGT) detection can be sensitive to the decoding choices of the source large language model (LLM). We observe that perturbing next-token logits or adjusting sampling temperature can reduce detection performance, providing a clear signal of detector vulnerability to decoding-time distribution changes. Building on this observation, we propose EASE (Entropy-Adaptive Distribution Shaping for Evasion), a training-free and detector-agnostic framework for evading AIGT detectors. EASE computes predictive entropy directly from the source LLM's next-token distribution and uses it to adapt both logit perturbation and sampling temperature, without detector feedback or model fine-tuning. Experiments across three source LLMs and multiple detectors demonstrate consistent reductions in detection performance, with negligible degradation in text quality and negligible inference overhead.