arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05741cs.CLcs.AI

一次响应,始终是响应:通过潜在提示恢复检测大语言模型生成的文本

Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration

Hongrui Bao, Yubing Ren, Jinhan You, Fang Fang, Shi Wang, Yanan Cao

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有零样本LLM生成文本检测器未考虑LLM训练过程的问题,提出无训练检测器EchoPrompt,通过潜在提示恢复实现零样本检测,在零样本检测器中达SOTA性能且稳健性强。

中文摘要 AI 辅助

大语言模型(LLMs)能够大规模生成流畅且有说服力的文本,这为错误信息传播、教育滥用和平台治理带来了日益增长的风险。这些担忧使得对机器生成文本的稳健检测愈发必要。近期的零样本检测器主要利用基于概率的统计差异,但它们未明确考虑LLMs的训练过程,这导致对独特的生成机制建模不足,限制了检测的稳健性。为解决该问题,我们提出EchoPrompt,一种基于潜在提示恢复的无训练检测器。我们的核心直觉是,机器生成的文本通常基于上游提示生成,而这种隐藏的依赖关系可通过前置一个统一的通用前缀部分重新激活。具体而言,EchoPrompt恢复了一个通用的助手-响应上下文,使用指令调优模型测量诱导的似然增益,针对对应的基础模型对其进行校准,并将得到的差异聚合成一个量化潜在提示依赖关系的分数。大量实验表明,EchoPrompt在零样本检测器中达到了最先进的性能,同时在具有挑战性的评估设置中保持了较强的稳健性。

英文摘要

Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, educational misuse, and platform governance. These concerns make robust detection of machine-generated text increasingly necessary. Recent zero-shot detectors mainly exploit probability-based statistical discrepancies, but they do not explicitly account for the training process of LLMs, which leaves a distinct generation mechanism insufficiently modeled and limits detection robustness. To address this issue, we propose EchoPrompt, a training-free detector based on latent prompt restoration. Our key intuition is that machine-generated text is typically produced conditioned on an upstream prompt, and this hidden dependency can be partially reactivated by prepending a unified generic prefix. Specifically, EchoPrompt restores a generic assistant-response context, measures the induced likelihood gain with an instruction-tuned model, calibrates it against the corresponding base model, and aggregates the resulting differences into a score that quantifies latent prompt dependency. Extensive experiments show that EchoPrompt achieves state-of-the-art performance among zero-shot detectors while maintaining strong robustness across challenging evaluation settings.

发表机构

  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
  • School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
  • College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑