arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越全局标量:结合令牌级统计与深度语义的对抗性AIGC文本检测

Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection

Peiming Li, Yifan Wang, Zhiyuan Hu, Shiyu Li, Zheng Wei, Yang Tang

arXiv 2608.28009首次发表:更新:

发表机构

Tencent BAC; School of Electronic and Computer Engineering, Peking University(腾讯BAC; 北京大学电子与计算机工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有AIGC文本检测方法在对抗场景中的缺陷,提出NeuroStat框架,结合令牌级统计与深度语义,在MOSAIC基准上实现鲁棒性提升,建立对抗性文本检测新基准。

AI 中文摘要

大型语言模型的快速发展催生了对鲁棒的机器生成文本检测技术的需求。现有范式通常分为两个独立分支:无训练方法依赖困惑度等全局统计标量,基于训练的方法则利用语义隐藏状态。两种方法在对抗场景中均存在根本性缺陷:全局标量作为有损压缩,会掩盖交错文本中的局部概率突发特性;而纯语义模型会过拟合特定指纹,仍易被欺骗。为揭示这些缺陷,我们推出MOSAIC——一个包含16000个样本、覆盖全粒度攻击谱的综合对抗基准。为应对这些挑战,我们提出NeuroStat,一个弥合统计与语义鸿沟的端到端框架。NeuroStat从单个因果语言模型主干中捕获未压缩的令牌级概率对数几率与深度语义隐藏状态,通过宏观状态残差调制融合这些异构信号,该调制利用全局不确定性指标自适应校准局部卷积特征;正交损失与对比损失进一步确保互补表示的学习。大量实验表明,相较于现有最优方法的严重性能下降,NeuroStat在MOSAIC基准上保持了出色的鲁棒性,为对抗性文本检测建立了新的标准。代码与MOSAIC基准可在该https网址获取。

英文摘要

The rapid evolution of large language models necessitates robust machine-generated text detection. Existing paradigms typically follow two isolated tracks. Training-free methods rely on global statistical scalars such as perplexity, while training-based methods utilize semantic hidden states. Both approaches exhibit fundamental vulnerabilities in adversarial scenarios. Global scalars act as lossy compressions that obscure local probabilistic burstiness in interleaved texts, whereas pure semantic models overfit to specific fingerprints and remain susceptible to spoofing. To expose these flaws, we introduce MOSAIC, a comprehensive adversarial benchmark comprising 16000 samples across a full-granularity attack spectrum. To address these challenges, we propose NeuroStat, an end-to-end framework bridging the statistical and semantic gap. NeuroStat captures uncompressed token-level probabilistic logits alongside deep semantic hidden states from a single causal language model backbone. We fuse these heterogeneous signals through Macro-State Residual Modulation, which adaptively calibrates local convolutional features using global uncertainty indicators. Orthogonal and contrastive losses further ensure the learning of complementary representations. Extensive experiments demonstrate that NeuroStat maintains exceptional robustness on MOSAIC compared to the severe degradation of state-of-the-art methods, establishing a new standard for adversarial text detection. Code and the MOSAIC benchmark are available at https://github.com/TencentBAC/NeuroStat.

CommentsAccepted by EMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑