arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DWT-Fusion:一种用于无训练大语言模型生成文本检测的基于信号的框架

DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection

Mehmet Batuhan Özdaş, Murat Osmanoğlu

arXiv 2607.22026首次发表:更新:

发表机构

Ankara University(安卡拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在零样本和无训练条件下检测大语言模型生成文本的挑战,提出基于离散小波分析的DWT-Fusion框架,通过多分辨率信号表示和校准引导投票融合检测文本,实验表明该方法能提供有效可解释信号,提升检测性能。

AI 中文摘要

在零样本和无训练条件下,检测大语言模型生成的文本仍然具有挑战性,特别是当检测器必须在数据集、领域和未见生成器之间进行泛化时。现有无训练方法利用语言模型统计信息作为检测信号,但通常通过总结整体模型行为的全局度量来表征文本,可能未充分利用令牌级可预测性中潜在的信息丰富的局部和多尺度变化。基于此,我们引入了DWT-Fusion,这是一个基于信号的无训练框架,用于使用代理因果语言模型生成的令牌级对数概率序列的离散小波分析来检测大语言模型生成的文本。该框架通过基于小波的多分辨率信号表示来分析这些序列,并从局部概率动态中导出检测信号。我们还评估了四种无训练投票变体,包括等权重硬投票、等权重软投票、校准加权硬投票和校准加权软投票,以在不训练监督元分类器的情况下组合多个小波配置。我们使用GPT-Neo-2.7B、GPT-J-6B、Falcon-7B和LLaMA-3-8B作为代理模型在HC3、M4和MAGE上评估了该框架。最佳单小波配置在HC3、M4和MAGE上的AUROC值分别为0.9872、0.8185和0.7138。通过校准加权投票,最佳集成变体进一步将AUROC提高到0.9919、0.8477和0.7471。这些发现表明,基于小波的多分辨率评分和校准引导的投票融合为无训练大语言模型生成文本检测提供了有效且可解释的信号。

英文摘要

Detecting LLM-generated text remains challenging under zero-shot and training-free conditions, especially when detectors must generalize across datasets, domains, and unseen generators. While existing training-free approaches exploit language-model statistics as detection signals, they typically characterize a text through global measures that summarize overall model behavior. Consequently, potentially informative local and multiscale variations in token-level predictability may remain underutilized. Motivated by this observation, we introduce DWT-Fusion, a training-free signal-based framework for detecting LLM-generated text using discrete wavelet analysis of token-level log-probability sequences produced by a proxy causal language model. The proposed framework analyzes these sequences through wavelet-based multiresolution signal representations and derives detection signals from localized probability dynamics. We further evaluate four training-free voting variants, including equal-weight hard voting, equal-weight soft voting, calibration-weighted hard voting, and calibration-weighted soft voting, to combine multiple wavelet configurations without training a supervised meta-classifier. We evaluate the framework on HC3, M4, and MAGE using GPT-Neo-2.7B, GPT-J-6B, Falcon-7B, and LLaMA-3-8B as proxy models. The best single wavelet configurations achieve AUROC values of 0.9872, 0.8185, and 0.7138 on HC3, M4, and MAGE, respectively. With calibration-weighted voting, the best ensemble variants further improve AUROC to 0.9919, 0.8477, and 0.7471. These findings show that DWT-based multiresolution scoring and calibration-guided voting fusion provide effective and interpretable signals for training-free LLM-generated text detection.

Comments40 pages, 4 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑