arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02177cs.CRcs.LG

WeaveMark:基于编码载荷扩散的鲁棒可扩展多比特大语言模型水印方案

WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading

Gang-Hyun Park, Ju-Hyeong Lee, Hee-Youl Kwak, Dae-Young Yun

首次发表
浏览论文内容

中文总结 AI 辅助

WeaveMark是一种基于编码载荷扩散的多比特LLM水印方案,通过多比特扩散、软判决纠错码等技术突破提取准确率、文本质量与载荷容量的权衡,在长消息、编辑文本及抗替换攻击场景下表现显著优于现有方法BiMark。

中文摘要 AI 辅助

大语言模型(LLM)的多比特水印技术可通过在生成文本中嵌入用户可识别的消息实现内容源追踪。现有方法在提取准确率、文本质量和载荷容量之间存在根本的权衡关系。我们提出WeaveMark,一种基于编码载荷扩散的鲁棒可扩展多比特LLM水印方案。WeaveMark通过每令牌多比特扩散提升载荷容量,通过软判决纠错码提升提取准确率,通过无偏多层重加权保留文本质量,进一步引入专用零比特层实现可靠的水印存在检测。实验显示该方案取得大幅提升,尤其针对长消息和编辑后文本:在200令牌的32比特消息上,WeaveMark的匹配率达89.8%,而BiMark仅为20.8%;在200令牌的16比特消息遭受10%替换攻击时,WeaveMark维持86.0%的匹配率,BiMark仅为30.7%,同时保留了文本质量。我们的代码可在此URL获取。

英文摘要

Multi-bit watermarking for large language models enables content source tracing by embedding user-identifiable messages into generated text. Existing methods face a fundamental trade-off among extraction accuracy, text quality, and payload capacity. We propose WeaveMark, a robust and scalable multi-bit LLM watermarking scheme based on coded payload spreading. WeaveMark shifts this trade-off frontier by improving payload capacity through multi-bit-per-token spreading (weaving), improving extraction accuracy through soft-decision error-correcting codes, and preserving text quality through unbiased multilayer reweighting. It further introduces dedicated zero-bit layers for reliable watermark presence detection. Extensive experiments demonstrate substantial gains in extraction performance, especially for long messages and edited text, without degrading text quality. WeaveMark achieves an 89.8% match rate for 32-bit messages at 200 tokens, compared with 20.8% for BiMark. Under 10% substitution attacks on 16-bit messages at 200 tokens, it maintains 86.0% versus 30.7%. Code is available at https://anonymous.4open.science/r/WeaveMark-ED6F.

发表机构

  • University of Ulsan(蔚山大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑