arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

T5-CSBoost:抗对抗扰动的语言模型指纹识别

T5-CSBoost: Adversarial Perturbation Resistant LLM Fingerprinting

Gayan K. Kulatilleke, Mahsa Baktashmotlagh, Siamak Layeghy, Marius Portmann

arXiv 2607.14113首次发表:更新:

AI 中文总结

研究针对AIGT检测器在多种干扰下准确性下降的问题,提出T5-CSBoost,通过引入辅助损失鼓励学习抗扰动风格表示,在多基准测试中达先进水平,对高强度对抗扰动鲁棒性增强,证明对比学习规范风格嵌入可构建更强大指纹识别系统。

AI 中文摘要

虽然许多人工智能生成文本(AIGT)检测器在干净输入上表现出色,但在轻度释义、单词替换、字符编辑和分布转移下,其准确性会显著下降。我们提出了T5对比风格增强分类器(T5-CSBoost),它是T5-Sentinel框架的扩展,在进行源归因时保持原来的下一个token预测目标,同时在解码器嵌入上引入基于边际的辅助三元组损失。这种对比风格正则化鼓励学习紧凑、抗扰动的风格表示。T5-CSBoost在OpenLLMText和HC3 AIGT基准测试中实现了多类源归因和二进制人类与语言模型检测的最先进水平。更重要的是,T5-CSBoost在高达90%强度的单词和字符级对抗扰动下表现出更强的鲁棒性。我们的结果表明,通过对比学习明确规范风格嵌入是在实际对抗环境中构建更强大语言模型指纹识别系统的实用有效策略。

英文摘要

While many AI-generated text (AIGT) detectors achieve strong performance on clean inputs, their accuracy degrades significantly under light paraphrasing, word substitutions, character edits, and distribution shifts. We present T5 Contrastive Style Boosted Classifier (T5-CSBoost), an extension to the T5-Sentinel framework that keeps the original next-token prediction objective for source attribution while introducing an auxiliary margin-based triplet loss over decoder embeddings. This contrastive style regularization encourages the learning of compact, perturbation-resistant stylistic representations, offering a lightweight yet effective alternative to prior approaches that rely on architectural modifications, adversarial training, or complex multi-task objectives without altering the underlying T5-small backbone. T5-CSBoost achieves state-of-the-art multiclass source attribution and binary human-vs-LLM detection on OpenLLMText and HC3 AIGT benchmarks. More importantly, T5-CSBoost demonstrates enhanced robustness to word and character level adversarial perturbations of up to 90% intensity, achieving state-of-the-art on the challenging MAGE/Deepfake stress-test suite, including unseen models, unseen domains, and extreme paraphrasing scenarios. Our results highlight that explicitly regularizing stylistic embeddings via contrastive learning is a practical and effective strategy for building more robust LLM fingerprinting systems in real-world adversarial settings.

Comments14 pages, 3 datasets, code and data will be provided

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑