arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00883cs.CLcs.CR

DeBERTa-ConPara:攻击感知且部署现实的AI生成文本检测

DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text

Mohamed Mady, Yupei Li, Johannes Reschke, Björn W. Schuller

首次发表
浏览论文内容

中文总结 AI 辅助

针对部署场景下AI生成文本检测的分布偏移与对抗扰动问题,提出DeBERTa-ConPara,采用原始训练加推理归一化预处理,在RAID隐藏测试上达到99.61% AUROC,显著提升同形字和零宽度空格攻击检测率。

中文摘要 AI 辅助

在部署条件下对AI生成文本进行稳健检测具有挑战性:跨领域和生成器的分布偏移、输入表面的对抗性扰动以及缺乏目标域标签进行阈值校准,都会削弱在域内表现良好的检测器。我们提出DeBERTa-ConPara,一种面向部署的检测器,结合了攻击感知的Unicode预处理和基于HC3 Plus、M4、MAGE和RAID训练的上下文Transformer编码器。我们的核心发现是,预处理的作用方向取决于其应用位置:对训练语料进行归一化会使其去重,将RAID中35.4%的行折叠为其干净副本,并删除对抗性监督,而在推理时进行归一化则是一种有效的防御。一个独立改变两个位置的因子实验确定,原始训练加归一化推理是最佳配置,在官方RAID隐藏测试上达到99.61%的AUROC、99.01%的TPR@5% FPR和96.57%的TPR@1% FPR,同时在固定阈值下,在HC3 Plus和MAGE上平均平衡准确率为93.14%。增益仅限于十二种攻击类别中的两类:同形字和零宽度空格插入分别从11.05%和1.12%提升至96.98%。相同的特征在另一种架构的零样本检测器中重现,表明该效应属于攻击而非我们的模型。我们还报告了两个负面结果:通过释义和监督对比学习(ConPara)进行的语义不变增强并未改善最佳配置,而手工特征融合分支在分布内无效且在分布外有害。

英文摘要

Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain. We present DeBERTa-ConPara, a deployment-oriented detector combining attack-aware Unicode preprocessing with a contextual transformer encoder trained over HC3 Plus, M4, MAGE and RAID. Our central finding is that preprocessing acts in opposite directions depending on where it is applied: normalising the training corpus deduplicates it, collapsing 35.4% of RAID rows into copies of their clean siblings and deleting the adversarial supervision, whereas normalising at inference is an effective defence. A factorial varying the two placements independently identifies raw training with normalised inference as the best configuration, reaching 99.61% AUROC, 99.01% TPR@5% FPR and 96.57% TPR@1% FPR on the official RAID hidden test, alongside 93.14% average balanced accuracy across HC3 Plus and MAGE under a fixed threshold. The gain is confined to two of twelve attack classes: homoglyph and zero-width-space insertion rise from 11.05% and 1.12% to 96.98%. The same signature reproduces in a zero-shot detector of different architecture, showing the effect belongs to the attacks rather than to our model. We additionally report two negative results: semantic-invariance augmentation through paraphrasing and supervised contrastive learning (ConPara) does not improve the best configuration, and the handcrafted feature-fusion branch is inert in distribution and harmful outside it.

发表机构

  • Technical University of Munich(慕尼黑工业大学)
  • OTH Regensburg(雷根斯堡应用技术大学)
  • Imperial College London(伦敦帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑