arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16681cs.CR

MarkSec:针对LLM水印对抗攻击的能力感知评估

MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Kairong Li, Zhikun Zhang, Xiao Ren, Yunjun Gao

AI总结:

针对LLM水印面临窃取、擦除和伪造攻击且评估缺乏统一协议的问题,提出MarkSec框架,引入质量约束的攻击成功率指标,发现攻击效果受文本质量约束和通用性影响。

AI中文摘要:

LLM水印技术有助于追踪生成文本的来源,但面临窃取攻击(恢复水印信息)、擦除攻击(移除水印信号)以及伪造攻击(伪造被认定为含水印的文本)的威胁。这些攻击通常被孤立研究,导致它们之间的联系不明确。评估也常常缺乏共享的检测器校准、指标定义和报告协议。此外,将攻击成功率和文本质量分开衡量,使得难以识别既有效又保持质量的攻击。我们提出MarkSec,一个统一分析窃取、擦除和伪造攻击的通用框架。我们在统一的报告协议下评估攻击,并引入一个质量约束的攻击成功率指标,以联合评估有效性和文本质量。在代表性水印家族、攻击、LLM和数据集上的实验揭示了三个发现。第一,仅凭水印移除效果看似最强的攻击,在成功率同时要求可接受的文本质量时,可能落后于通用改写。第二,通用改写仍然是跨水印家族的强基线,而它相对于其他擦除器的优势因家族而异。第三,在一个水印家族的案例研究中,当要求文本质量时,基于窃取的擦除器往往表现不如最佳的通用擦除基线。这些结果表明,明显的攻击赢家取决于文本质量约束、攻击通用性和能力假设。

英文摘要:

LLM watermarking helps trace the origin of generated text, but faces stealing attacks that recover watermark information, scrubbing attacks that remove watermark signals, and spoofing attacks that forge text accepted as watermarked. These attacks are often studied in isolation, leaving their connections unclear. Evaluations also often lack shared detector calibration, metric definitions, and reporting protocols. Moreover, measuring attack success and text quality separately makes it difficult to identify attacks that are both effective and quality-preserving. We propose MarkSec, a general framework that unifies analyses of stealing, scrubbing, and spoofing. We evaluate attacks under a common reporting protocol and introduce a quality-constrained attack success metric to assess effectiveness and text quality jointly. Experiments across representative watermark families, attacks, LLMs, and datasets reveal three findings. First, attacks that appear strongest by watermark removal alone can fall behind general rewriting when success also requires acceptable text quality. Second, general rewriting remains a strong baseline across watermark families, while its advantage over other scrubbers varies by family. Third, in a case study of one watermark family, stealing-based scrubbers often underperform the best general-scrubbing baselines when text quality is required. These results show that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions.

补充信息

↑