arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25944cs.CL

揭示无训练大语言模型文本检测中的频谱机制

Unveiling Spectral Mechanisms in Training-Free LLM Text Detection

  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Haitong Luo, Xuying Meng, Weiyao Zhang, Wenji Zou, Shengfeng Lou, Xuefeng Jiang, Chungang Lin, Yujun Zhang

AI总结:

该研究针对无训练LLM文本检测中遗漏人类写作信号波动的问题,从理论和实证角度分析频谱检测机制,明确其适用场景,为多维检测器设计提供指导。

AI中文摘要:

大语言模型(LLM)的快速发展使得区分人类写作与机器生成文本愈发困难。无训练检测提供了一种可扩展的解决方案,但常见的基于置信度的指标主要测量平均词元概率,往往遗漏了人类写作特有的信号波动,我们将其称为“生成活力”。频谱分析为捕捉这种活力提供了途径,但其机制和实际边界仍未得到充分探索。本文从理论和实证角度分析了频谱检测,将频谱能量与代理对数概率轨迹的方差关联起来,解释了人类更广泛的词元选择如何产生频域指标所利用的波动。我们进一步表明,该信号的强度取决于文本长度和采样范围:频谱证据在长、连续、受约束的生成中最为清晰,而在短、碎片化、混合及编辑场景中则需要互补的置信度和波动视图。这些发现明确了频域检测的适用场景,并为未来多维检测器的设计提供了指导。

英文摘要:

The rapid advancement of Large Language Models (LLMs) makes it increasingly difficult to distinguish human writing from machine-generated text. Training-free detection offers a scalable solution, yet common confidence-based metrics mainly measure average token probabilities and often miss the signal fluctuations that characterize human writing, which we call "generative vitality". Spectral analysis offers a way to capture this vitality, but its mechanism and practical boundaries remain underexplored. In this paper, we analyze spectral detection from both theoretical and empirical perspectives. We connect spectral energy to variance in proxy log-probability trajectories and explain how broader human token choices create the fluctuations used by frequency-domain indicators. We further show that the strength of this signal depends on text length and sampling range: spectral evidence is clearest for long, continuous, constrained generation, while short, fragmented, mixed, and edited settings require complementary confidence and fluctuation views. These findings clarify when frequency-domain detection works and provide guidance for future multi-dimensional detector design.

↑