arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型指纹识别需要重新思考水印教师

Language Model Fingerprinting Requires Rethinking Watermark Teachers

Jeongyeon Hwang, Anshul Nasery, Sewoong Oh, Jungseul Ok

arXiv 2610.04169首次发表:更新:

发表机构

Pohang University of Science and Technology; University of Washington(浦项科技大学; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文重新审视LLM指纹识别中的水印教师选择,提出近并列限制方法,通过top-1相对logit差距放置稀疏水印信号,在多个模型上改善了检测与质量的权衡。

AI 中文摘要

通过水印蒸馏实现的LLM指纹识别,将统计水印信号嵌入模型权重,使模型所有者能够在黑盒API背后识别其模型。重新审视近期的一个协议,我们发现其效用评估低估了开放生成中文本质量的下降,从而偏向于过强的水印教师。减弱水印可提高文本质量,但会牺牲可检测性。为超越这一权衡,我们重新思考为验证生成文本而设计的文本水印是否适合作为模型指纹识别的蒸馏教师。此类水印通常设计为能从单个输出中保持可检测性,限制了水印信号可以多稀疏。相比之下,指纹验证可以跨查询聚合信号,使更稀疏的水印信号变得可行。这引出一个关键问题:稀疏信号应放置在哪里?我们通过词元惊讶度分析信号放置,并表明即使在相当的水印强度下,不同放置可以针对基础模型下具有不同合理性的词元。这促使我们提出近并列限制,即使用top-1相对logit差距将水印偏差限制在接近基础模型顶部预测的词元上。在多个模型上,近并列在部署变化下改善了检测-质量前沿,跨查询预算保持更高的文本质量,并在与现有水印方案结合时进一步改进它们。

英文摘要

LLM fingerprinting via watermark distillation embeds a statistical watermark signal into model weights, enabling model owners to identify their models behind black-box APIs. Revisiting a recent protocol, we find that its utility evaluation understates text quality degradation in open-ended generation, favoring overly strong watermark teachers. Weakening the watermark improves text quality but sacrifices detectability. To move beyond this trade-off, we rethink whether text watermarks designed for verifying generated text are suitable distillation teachers for model fingerprinting. Such watermarks are typically designed to remain detectable from an individual output, limiting how sparse the watermark signal can be. In contrast, fingerprint verification can aggregate signal across queries, making sparser watermark signals viable. This raises a key question: where should the sparse signal be placed? We analyze signal placement through token surprisal and show that, even at comparable watermark strength, different placements can target tokens with different plausibility under the base model. This motivates near-tie restriction, which uses top-1-relative logit gaps to restrict the watermark bias to tokens close to the base model's top prediction. Across multiple models, near-tie improves detection--quality frontiers under deployment changes, preserves higher text quality across query budgets, and further improves existing watermarking schemes when combined with them.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑