arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeBERTa-Sentinel:实现透明且可信的AI生成文本检测

DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text

Muhammad Yousaf Rehman, Muhammad Islam

arXiv 2608.01046首次发表:更新:

发表机构

University of Hertfordshire; James Cook University(赫特福德大学; 詹姆斯库克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DeBERTa-Sentinel是基于DeBERTa-v3的透明AI生成文本检测框架,在GLC-AIText数据集上实现优异检测性能,可公开token级解释以支持利益相关者审计。

AI 中文摘要

大型语言模型(LLMs)在网络上的快速传播引发了人们对错误信息、学术诚信、自动内容操纵以及弱势网络社区风险的担忧。现有的基于Transformer的检测器,如GPT-Sentinel,虽有应用前景,但难以泛化到不同模型的输出和 paraphrasing 攻击,限制了其在构建可信网络生态系统中的作用。本研究提出DeBERTa-Sentinel,这是一种负责任的AI生成文本检测框架,利用DeBERTa-v3的解耦注意力机制捕捉合成内容中的细微结构异常。其核心设计原则是透明度:与黑箱式商业检测器不同,DeBERTa-Sentinel会公开其决策的 token 级解释,使受影响的利益相关者——记者、教育工作者以及平台信任与安全团队——能够对检测结果进行审计、质疑和情境化分析。使用包含28057个人类及LLM生成样本(来自GPT、LLaMA和Claude)的GLC-AIText数据集,按60-20-20划分,DeBERTa-Sentinel的验证准确率达98.21%,超过了NeurIPS 2025的RoBERTa-Sentinel基线,测试准确率为97.53%,精确率为95.89%,召回率为99.33%,ROC-AUC为99.53%,同时保持0.665%的假阴性率。该模型的可解释性揭示了与合成文本相关的语言标记,如学术措辞和正式过渡,直接支持利益相关者对可验证、可审计的内容真实性决策的需求。通过推进减少偏差、增强可解释性的负责任检测方法,DeBERTa-Sentinel推动了可信、合乎道德且以人类为中心的AI系统发展。代码和数据可在指定URL获取。

英文摘要

The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation, and risks to vulnerable online communities. Existing transformer-based detectors, such as GPT-Sentinel, show promise but struggle to generalize to diverse model outputs and paraphrasing attacks, limiting their role in building trustworthy web ecosystems. This work introduces DeBERTa-Sentinel, a responsible AI-generated text detection framework leveraging DeBERTa-v3's disentangled attention to capture subtle structural irregularities in synthetic content. A central design principle is transparency: unlike black-box commercial detectors, DeBERTa-Sentinel exposes token-level explanations of its decisions, enabling affected stakeholders journalists, educators, and platform trust and safety teams to audit, challenge, and contextualize detection outcomes. Using the GLC-AIText dataset of 28,057 human and LLM-generated samples (GPT, LLaMA, and Claude) with a 60-20-20 split, DeBERTa-Sentinel achieves 98.21\% validation accuracy and surpasses the RoBERTa-Sentinel baseline from NeurIPS 2025, achieving 97.53\% test accuracy, 95.89\% precision, 99.33\% recall, and 99.53\% ROC-AUC, and maintaining a 0.665\% false negative rate. The model's interpretability reveals linguistic markers such as academic phrasing and formal transitions associated with synthetic text, directly supporting stakeholder needs for verifiable, auditable content-authenticity decisions. By advancing responsible detection methods that reduce bias and enhance explainability, DeBERTa-Sentinel promotes trustworthy, ethical, and human-centric AI systems. Code and data are available at https://github.com/Galileo-Galili/HUMAN-VS-AI-TEXT-DETECTION.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑