arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成模型的水印取证:信息论视角

Watermark Forensics for Generative Models: An Information-Theoretic Perspective

Xiaoyu Li, Zheng Gao, Xiaoyan Feng, Jiaojiao Jiang, Yulei Sui, Jiankun Hu

arXiv 2607.13003首次发表:更新:

发表机构

University of New South Wales; Griffith University(新南威尔士大学; 格里菲斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

从信息论视角研究生成模型水印取证,通过信息轮廓确定取证阶梯各级成本,给出多用户归属和提取有效载荷的熵率定律,实验验证了相关预测常数,揭示了水印取证中的实际差距。

AI 中文摘要

生成模型输出中的水印通常仅用于判断文本是否由机器生成。实际上,同一水印还有更多用途,如将文本归属于生成用户、提取隐藏有效载荷或定位编辑后仍保留的部分,这些构成了一个取证阶梯。我们研究了在样本长度\(n\)下,阶梯的每一级成本。通过信息轮廓\(\nu(t)=I(S;X_t\mid X_{<t})\)来组织答案,其总质量用于归属和提取,质量分布用于定位,检测则基于标记与未标记分布的距离。文献中的两种质量模型是限制该轮廓的两种不可比方式。我们的主要定理确定了阶梯的熵列。对于统计无失真方案,将文本归属于\(N\)个用户之一,在每个熵率为\(h\)的平稳遍历源上,成本为\(\Theta(\log N/h)\)个令牌,这是多用户归属的首个精确熵率定律。提取\(\ell\)位有效载荷的成本为\(\Theta(\ell/h)\)。存在两个实际差距,而非建模伪像:一个\(\Theta(\log N)\)令牌的窗口,其中文本可证明是机器生成但无法归属,以及一个足迹分辨率不确定性原理。在GPT - 2、Pythia - 得410M和Qwen2.5上的实验验证了预测常数。

英文摘要

A watermark in a generative model's output is usually asked only whether a text is machine-made. The same mark can do more: attribute it to the user who produced it, extract a hidden payload, or localize the part that survives editing. These form a forensic ladder, and we ask what each rung costs in the sample length $n$. One object organizes the answers. Let $S$ be the secret the mark carries (a user's identity or payload), and let the information profile $ν(t)=I(S;X_t\mid X_{<t})$ record how much the $t$-th token reveals about $S$ given the earlier ones. Its total mass pays for attribution and extraction; how that mass is spread pays for localization; and detection alone is paid for not by information but by presence, the distance from the marked to the unmarked distribution. The literature's two quality models, a mark subtle on every token and one that stamps a few tokens loudly, are two incomparable ways of capping this profile. Our main theorem settles the ladder's entropy column. For statistically distortion-free schemes, attributing a text to one of $N$ users costs $Θ(\log N/h)$ tokens over every stationary-ergodic source of entropy rate $h$, sharp to a $(1+o(1))$ factor: to our knowledge the first tight entropy-rate law for multi-user attribution (via exact alignment). The natural collision-counting analysis overcharges without bound; only a decoder thresholding each candidate by its own realized surprisal attains the rate while almost never implicating an innocent user. A matching converse makes the law two-sided, and extraction of an $\ell$-bit payload costs $Θ(\ell/h)$. Two gaps are real, not modeling artifacts: a $Θ(\log N)$-token window in which a text is provably machine-made yet unattributable, and a footprint-resolution uncertainty principle. Experiments on GPT-2, Pythia-410M, and Qwen2.5 recover the predicted constants.

CommentsThe abstract has been shortened to comply with arXiv's length limit

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑