发表机构
Johannes Kepler University; CONICET - UNICEN; Aalborg University(约翰开普勒林茨大学; 国家科学与技术研究理事会 - 国立中部大学; 奥尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 SHADE 框架,通过 GRPO 微调 LLaMA 模型以规避 AI 文本检测器,实验显示全参数微调达到 98.5% 的代理规避率,但对抗单一检测器的优化难以泛化到未见分类器。
AI 中文摘要
大型语言模型(LLMs)能够生成流畅且连贯的文本,这些文本越来越难以与人类写作区分开来,这推动了自动 AI 生成文本检测器的发展。然而,这些检测器在对抗性生成条件下的鲁棒性仍不确定。本文提出了 SHADE(通过对抗性检测器规避进行随机类人生成),这是一种强化学习框架,将检测器规避表述为策略优化问题。SHADE 不采用事后扰动或基于提示的重写,而是使用组相对策略优化(GRPO)对经过指令微调的 LLaMA 模型进行微调,并利用基于 PAN 2025 mdok 系统的代理检测器的反馈。我们的实验表明,使用较小的 KL 正则化惩罚进行全参数微调可实现 $98.5\%$ 的代理规避率,而基础模型仅为 $1.5\%$,同时基于 LoRA 的适配在正则化下效果明显较差。语言学分析显示,成功的规避与更短、更简单且词汇多样性更低的输出相关,这表明高检测器规避率并不一定对应于更像人类的写作。在官方的 Voight-Kampff 竞赛设置中,我们的提交排名第六和第七,表明针对单一代理检测器的优化仅部分迁移到未见过的评估分类器。这些结果凸显了强化学习在对抗性 AI 文本生成中的潜力和局限性,并促使为 AI 生成文本检测开发更鲁棒的多检测器评估协议。
英文摘要
Large language models (LLMs) can generate fluent and coherent text that is increasingly difficult to distinguish from human writing, motivating the development of automatic AI-generated text detectors. However, the robustness of such detectors under adversarial generation remains uncertain. This paper presents SHADE (Stochastic Human-like generation via Adversarial Detector Evasion), a reinforcement learning framework that formulates detector evasion as a policy optimization problem. Instead of applying post-hoc perturbations or prompting-based rewriting, SHADE fine-tunes an instruction-tuned LLaMA model with Group Relative Policy Optimization (GRPO), using feedback from a surrogate detector based on the PAN 2025 mdok system. Our experiments show that full fine-tuning with a small KL regularization penalty achieves $98.5\%$ surrogate evasion, compared to $1.5\%$ for the base model, while LoRA-based adaptation is substantially less effective under regularization. Linguistic analysis reveals that successful evasion is associated with shorter, simpler, and less lexically diverse outputs, suggesting that high detector evasion does not necessarily correspond to more human-like writing. In the official Voight-Kampff competition setting, our submissions ranked sixth and seventh, indicating that optimization against a single surrogate detector only partially transfers to unseen evaluation classifiers. These results highlight both the potential and limitations of reinforcement learning for adversarial AI-text generation and motivate more robust, multi-detector evaluation protocols for AI-generated text detection.
CommentsAccepted for Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2026)