arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29030cs.AI

通过自蒸馏学习遵循上下文水印指令

Learning to Follow In-Context Watermark Instructions via Self-Distillation

Yepeng Liu, Tianyi Chen, Xuandong Zhao, Dawn Song, Yuheng Bu

首次发表
浏览论文内容

中文总结 AI 辅助

针对当前LLM无法同时遵循上下文水印指令并保持答案质量的问题,提出带logits扰动的自蒸馏(SDLP)与强化学习结合的两阶段训练方法,在Qwen3-14B和GPT-OSS-20B上显著提升水印可检测性且维持响应质量。

中文摘要 AI 辅助

上下文水印(ICW)会在查询前添加一条指令,要求模型在响应中嵌入具有统计可检测性的信号,从而为大型语言模型(LLM)配备了第三方无需访问模型内部即可调用的水印接口。其可靠性取决于LLM遵循指令的同时不降低答案质量,但当前LLM的遵循效果尚未被测量。我们推出ICWBench,这是一个包含三类可验证ICW指令的基准,每类指令均在可检测性和答案质量两方面评分。对14个前沿专有和开源LLM的评估显示,所有被评估的LLM均未在三类指令中同时实现两个目标。为解决此问题,我们提出一种自包含的两阶段训练方法,无需更强模型的蒸馏、人工标注或现有的ICW IF能力。第一阶段是带logits扰动的自蒸馏(SDLP),使用同一基础LLM作为教师和学生:与指令等效的解码时logits扰动使教师遵循ICW指令,学生则被训练以匹配教师的输出分布。第二阶段使用自动验证器作为奖励应用强化学习。将该方法应用于Qwen3-14B和GPT-OSS-20B时,我们的方法将三类ICW指令的平均TPR@1%FPR分别从0.100提升至0.974和从0.337提升至0.968,同时在困惑度评估和LLM-as-a-Judge下保持高响应质量。

英文摘要

In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarking interface that third parties can invoke without access to model internals. Its reliability hinges on the LLM following the instruction without degrading answer quality, yet how well current LLMs do so has not been measured. We introduce $\mathsf{ICWBench}$, a benchmark of three verifiable ICW instruction families, each scored on both detectability and answer quality. Evaluating 14 frontier proprietary and open-source LLMs, we find that none of the evaluated LLMs achieves both objectives across all three families. To address this, we propose a self-contained two-stage training method, requiring no distillation from a stronger model, no manual annotation, and no pre-existing ICW IF ability. The first stage, self-distillation with logits perturbation (SDLP), uses the same base LLM as both teacher and student: an instruction-equivalent decoding-time logits perturbation makes the teacher follow the ICW instruction, and the student is trained to match the teacher's output distribution. The second stage applies reinforcement learning with the automatic verifier as the reward. Applied to Qwen3-14B and GPT-OSS-20B, our method raises average TPR@$1\%$FPR across three ICW instructions from $0.100$ to $0.974$ and from $0.337$ to $0.968$, respectively, while maintaining high response quality under both perplexity evaluation and LLM-as-a-Judge.

发表机构

  • UC Santa Barbara(加州大学圣巴巴拉分校)
  • UC San Diego(加州大学圣迭戈分校)
  • UC Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

↑