TANGO:在Token对中为掩码扩散语言模型添加水印
TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
AI总结:
TANGO提出一种基于token对颜色关联的水印方法,用于掩码扩散语言模型,通过动态颜色偏好保持频率平衡,实现高检测率并抵抗频率伪造攻击。
AI中文摘要:
掩码扩散语言模型以并行且无固定顺序的方式填充掩码位置。大多数实用的文本水印假设从左到右生成。它们将每个token与其之前的token关联,而在扩散模型中,这些token可能仍处于掩码状态。固定绿色列表不需要这样的上下文,但它会在每个位置偏向相同的token,因此这些token在水印文本中出现得更频繁。攻击者通过比较水印文本和未水印文本中的token频率,可以恢复该列表并伪造出提供商自己的检测器会接受的文本。我们提出了TANGO,一种针对掩码扩散语言模型的水印方法,它将每个新token与一个附近已解除掩码的token关联。一个密钥将词汇表划分为颜色类别,TANGO根据密钥和附近token的颜色,将新token偏向于由该颜色决定的类别。因此,水印被嵌入到token对中。由于偏好的颜色随位置变化,token频率与未水印文本相比,比固定绿色列表下更接近。检测仅需文本和密钥,且不假设任何解除掩码的顺序。在两个掩码扩散模型上,TANGO能检测出几乎所有未编辑的水印文本和大多数已编辑的文本,而伪造固定绿色列表的频率攻击对其无效。
英文摘要:
Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked. A fixed green list needs no such context, but it favors the same tokens at every position, so these tokens appear more often in watermarked text. An attacker who compares token frequencies in watermarked and unwatermarked text can recover the list and forge text that the provider's own detector accepts. We present TANGO, a watermark for masked-diffusion language models that keys each new token to a nearby token that is already unmasked. A secret key splits the vocabulary into color classes, and TANGO biases the new token toward a color determined by the key and the nearby token's color. The watermark is therefore embedded in pairs of tokens. Because the favored color changes from position to position, token frequencies stay much closer to those of unwatermarked text than under a fixed green list. Detection needs only the text and the key, and it does not assume any unmasking order. On two masked-diffusion models, TANGO detects nearly all unedited watermarked texts and most edited ones, and frequency attacks that forge the fixed green list fail against it.