arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向机器生成文本鲁棒水印检测的感知稳定性特征设计

Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text

Sina Mansouri, Mohit Marvania, Abolfazl Safikhani

arXiv 2608.18102首次发表:更新:

AI 中文总结

该研究针对机器生成文本水印检测在多次释义及短文本场景下性能骤降的问题,提出PSS检测框架,结合多类特征与稳定性得分,在多基准数据集和多模型测试中显著提升检测AUC,且通用分类器跨域泛化性强。

AI 中文摘要

大型语言模型(LLMs)的广泛应用加剧了区分人类文本与机器生成文本的原理性方法需求。水印技术提供了一条有前景的途径,但现有检测器在多次释义及应用于较短文本时性能会急剧下降。我们提出Pattern Stability Score(PSS,模式稳定性得分),这是一种利用局部统计特征和释义变体间稳定性动态的新型检测框架。具体而言,该方法将全局与局部z分数特征、游程模式的高阶统计量相结合,辅以自相关信号和基于释义深度计算的稳定性得分。在三个基准数据集(PG-19、CNN/DailyMail和WikiText)上使用多个LLMs(Llama-3-8B、Qwen2-7B)和释义器(Mistral-7B、Qwen2-7B、Gemma-7B)进行数值评估,系统地对最多八轮释义下的鲁棒性进行压力测试。与现有的z分数阈值基线及部分最先进的深度学习方法相比,我们的方法在不同令牌长度下将检测AUC(受试者工作特征曲线下面积)提升了10至15个百分点以上。此外,广泛的跨域实验表明,单个通用分类器无需重新训练即可在不同LLMs、释义器和文本域间泛化,即使所有组件与训练时不同,仍能保持87.8%以上的AUC。

英文摘要

The widespread adoption of large language models (LLMs) has intensified the demand for principled methods to distinguish human from machine-generated text. Watermarking provides a promising avenue, yet existing detectors exhibit sharp performance deterioration under multiple paraphrasing and when applied to shorter texts. We introduce Pattern Stability Score (PSS), a novel detection framework that leverages local statistical features and stability dynamics across paraphrased variants. Specifically, the proposed method combines global and local z-score features with higher-order statistics of run-length patterns, enriched by autocorrelation signals and stability scores computed over paraphrase depth. Numerical evaluations are performed on three benchmark datasets (PG-19, CNN/DailyMail, and WikiText) using multiple LLMs (Llama-3-8B, Qwen2-7B) and paraphrasers (Mistral-7B, Qwen2-7B, Gemma-7B), systematically stress-testing robustness under up to eight rounds of paraphrasing. Compared to prior z-score thresholding baselines and some state-of-the-art deep learning methods, our approach improves detection AUC (area under the receiver operating characteristic curve) by over 10-15 percentage points across different token lengths. Additionally, extensive cross-domain experiments demonstrate that a single universal classifier generalizes across different LLMs, paraphrasers, and text domains without retraining, maintaining above 87.8% AUC even when all components differ from training.

CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026), Seoul, South Korea. 20 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑