CertMark:具有认证解码的无失真多比特水印
CertMark: Distortion-Free Multi-Bit Watermarking with Certified Decoding
AI总结:
CertMark是一种无失真多比特水印方法,通过Gumbel-max采样保持分布,并提供带认证弃权(不执行)的解码器,在保证文本质量的同时可靠恢复消息,且模型感知解码器比特准确率更高。
AI中文摘要:
领先的多比特水印方法通过偏置模型的下一词概率来编码消息,从而在消息恢复和文本质量之间产生权衡。它们的解码器通常从累积的令牌级证据中返回得分最高的候选,而没有认证弃权(不执行)规则来限制输出错误消息的概率。我们引入了CertMark,一种具有认证解码的分布保持多比特水印。CertMark不修改概率,而是使用嵌入消息来种子化精确的Gumbel-max采样器,从而保持模型的原始采样分布。我们提出了两种可扩展的解码器:一种与模型无关的纯文本解码器,以及一种利用原始下一词分布以增强恢复的模型感知变体。两者都支持认证弃权(不执行),并对返回错误消息的概率提供数学界限。在文本补全、摘要生成和故事生成中,CertMark在可靠恢复多比特消息的同时,匹配了未加水印文本的困惑度。模型感知解码器进一步实现了比概率偏置基线更高的比特准确率。我们的代码可在https://this URL公开获取。
英文摘要:
Leading multi-bit watermarking methods for language models encode messages by biasing the model's next-token probabilities, creating a trade-off between message recovery and text quality. Their decoders typically return the highest-scoring candidate from accumulated token-level evidence, without a certified abstention rule that bounds the probability of outputting an incorrect message. We introduce CertMark, a distribution-preserving multi-bit watermark with certified decoding. Rather than modifying probabilities, CertMark uses the embedded message to seed an exact Gumbel-max sampler, thereby preserving the model's original sampling distribution. We propose two scalable decoders: a model-agnostic, text-only decoder and a model-aware variant that leverages the original next-token distributions for stronger recovery. Both support certified abstention with mathematical bounds on the probability of returning an incorrect message. Across text completion, summarization, and story generation, CertMark matches the perplexity of unwatermarked text while reliably recovering multi-bit messages. The model-aware decoder further achieves higher bit accuracy than probability-biasing baselines. Our code is publicly available at https://github.com/Batorskq/CertMark.