arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11848cs.CR

SemField:一种简单、线性、连续且鲁棒的语义水印

SemField: A Simple, Linear, Continuous, yet Robust Semantic Watermark

Varun Gumma, Navonil Majumdar, Soujanya Poria

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLMs水印易受改写攻击的问题,提出无训练语义水印SemField及其极性锁定变体,经评估其性能优于12种基线,在1% FPR下平均TPR达88.2%-90.4%,且保持良好自然度。

中文摘要 AI 辅助

大型语言模型(LLMs)的快速普及催生了对可靠水印技术的需求,以识别AI生成的文本并确保适当的归因。基于token的水印易受改写攻击影响,而语义水印的核心挑战在于将句子含义转化为稳定、校准良好的文档级信号。为此,我们提出SemField,一种简单的无训练语义水印,它将连续、线性信号直接嵌入句子嵌入空间。该方法使用共享密钥定义特定高斯方向,迭代评估候选句子并选择那些使累积文档聚合与目标方向对齐最大化的句子。我们还提出SemField-PL,一种极性锁定变体,它先从初始句子确定最优方向,再在生成过程中持续强化该方向。文档级聚合公式保证了对句子重排的完全不变性,并提供了针对句子插入、删除等结构篡改的理论边界。最后,通过在三个模型上的广泛评估,我们证明两种变体均优于12种近期基线方法;在1%假阳性率(FPR)下,两种变体在干净检测和四种改写攻击场景下的平均真阳性率(TPR)达到88.2%至90.4%,同时保持与人类生成内容相当的困惑度和自然度。我们在该https URL开源了实现代码。

英文摘要

The rapid proliferation of Large Language Models (LLMs) necessitates reliable watermarking techniques to identify AI-generated text and ensure appropriate attribution. While token-based watermarks are vulnerable to paraphrasing, a central challenge for semantic watermarking is to turn sentence meanings into a stable, well-calibrated document-level signal. To this end, we introduce SemField, a simple, training-free semantic watermark that embeds a continuous, linear signal directly into the sentence embedding space. Using a shared secret key to define a specific Gaussian direction, SemField iteratively evaluates candidate sentences and selects those that maximize the alignment of the cumulative document aggregate with this targeted direction. We also propose SemField-PL, a polarity-locked variant that first determines the optimal orientation from the initial sentence and continuously reinforces it throughout the generation process. The document-level aggregated formulation guarantees exact invariance to sentence reordering and provides theoretical bounds against structural tampering, such as sentence insertion and deletion. Lastly, with extensive evaluations across three models, we demonstrate that both variants outperform 12 recent baselines. Across clean detection and four paraphrasing attacks, both variants achieve a mean True Positive Rate (TPR) of 88.2% to 90.4% at a 1% False Positive Rate (FPR), all while maintaining comparable perplexity and naturalness as human-generated content. We open-source our implementation at https://github.com/declare-lab/SemField

发表机构

  • DeCLaRe Lab, Nanyang Technological University(南洋理工大学 DeCLaRe 实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑