发表机构
Chungnam National University; Nara Institute of Science and Technology (NAIST); Institute of Science Tokyo(忠南大学; 奈良先端科学技术大学院大学; 东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有文本水印方法降低生成质量的问题,提出基于三词汇分区和模式交替嵌入的TripPattern框架,在保持质量的同时实现稳健可检测的水印。
AI 中文摘要
文本水印技术因识别机器生成文本和减轻大型语言模型(LLM)风险而受到广泛关注。现有方法通常将LLM的词汇表划分为绿色和红色标记,但鼓励生成绿色标记可能会降低文本质量和自然度。为解决此问题,我们提出TripPattern,一个将文本水印表述为基于模式的匹配任务的水印框架,该框架使用三个词汇分区。TripPattern将词汇表划分为一个中性组和两个模式组。在生成过程中,模型在两个模式组之间交替选择标记以嵌入可检测的模式,而中性标记则被独立选择以提高灵活性并保持自然度。在检测方面,TripPattern使用基于模式的统计检验,通过测量相邻标记在模式组之间交替的频率来提供可解释的p值。理论分析和在四个多语言数据集上的实证评估表明,TripPattern在保持LLM生成质量的同时实现了稳健的水印可检测性。
英文摘要
Text watermarking techniques have gained significant attention for identifying machine-generated text and mitigating risks from large language models (LLMs). Existing methods typically divide an LLM's vocabulary into green and red tokens, but encouraging generation toward green tokens can reduce text quality and naturalness. To address this, we propose TripPattern, a watermarking framework that formulates text watermarking as a pattern-based matching task using three vocabulary partitions. TripPattern divides the vocabulary into one neutral group and two pattern groups. During generation, the model alternates token selection between the two pattern groups to embed detectable patterns, while neutral tokens are selected independently to improve flexibility and preserve naturalness. For detection, TripPattern uses pattern-based statistical tests that provide interpretable p-values by measuring how often adjacent tokens alternate between the pattern groups. Theoretical analysis and empirical evaluations on four multilingual datasets show that TripPattern maintains LLM generation quality while achieving robust watermark detectability.
CommentsAccepted to Findings of AACL-IJCNLP 2026. 16 pages, 4 figures, 10 tables