arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义行为水印:面向LLM智能体的释义鲁棒且防伪造溯源

Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents

Suxin Ji, Hungtao Wan, Shaoxuan Chen, An Zhang

arXiv 2610.08668首次发表:更新:

发表机构

University of Pennsylvania; University of Massachusetts Amherst; Chongqing University(宾夕法尼亚大学; 马萨诸塞大学阿默斯特分校; 重庆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出语义行为水印(SBW),通过语义动作聚类和带密钥抗碰撞分箱实现LLM智能体行为溯源,在释义重写下显著提升检测鲁棒性,并将自适应伪造降至假阳性下限,但链式重放攻击仍开放。

AI 中文摘要

行为水印将所有者标识嵌入LLM智能体的高层动作选择中,在不触碰输出令牌的情况下实现溯源。先前的智能体水印在两个方面存在缺陷。首先,所有三种先前方案都将水印绑定到精确的动作符号上,因此重命名工具会导致解码失同步,即使观察结果未被改动;在AgentMark自身的鲁棒性测试中,仅对观察结果进行释义就会使比特恢复率降至16.8%。其次,每种先前的智能体水印都只研究移除问题:没有人探究对手能否伪造一条可验证为他人所有的轨迹,而这个问题对于文本水印已有肯定答案(Jovanović等人,2024)。我们提出语义行为水印(SBW):在历史条件约束下对语义动作聚类进行水印,将公共聚类桶替换为带密钥的抗碰撞分箱,其新桶分配在随机预言机模型下被证明是不可预测的。在五个智能体模型(3B-14B,四个供应商)和三种编码器上,排序在两个基准上均成立:在ToolBench上(每个模型600条轨迹),重写下的检测率为聚类级别的0.49-0.66,对比精确符号的0.05-0.17(在置换校准的1%假阳性率下),选择一致性为72-83%,对比logit偏置的22-27%;在ALFWorld上(每个模型100个回合),检测率为0.92-0.97,对比0.00-0.01。带密钥分箱将自适应伪造从100%降至主操作点(bge,r=64)的假阳性下限。我们还标明了该保证未覆盖的边界:当对手复制受害者自身的步骤时,混洗拼接被中和(在Qwen2.5-3B上为0.000),但链式重放在五个模型上仍为0.76-0.98,作为开放问题报告。释义鲁棒性大约消耗每步水印容量的一半。代码可在https URL获取。

英文摘要

Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched; in AgentMark's own robustness test, paraphrasing the observation alone drops bit-recovery to 16.8%. Second, every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else's, a question answered affirmatively for text watermarks (Jovanović et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over semantic action clusters under history conditioning, with the public-cluster bin replaced by keyed collision-resistant binning whose fresh-bucket assignment is provably unpredictable in the random-oracle model. Across five agent models (3B-14B, four vendors) and three encoders the ordering holds on both benchmarks: on ToolBench (600 trajectories per model) detection under rewriting is 0.49-0.66 for cluster-level versus 0.05-0.17 for exact-symbol at a permutation-calibrated 1% FPR, at 72-83% choice agreement against 22-27% for logit biasing; on ALFWorld (100 episodes per model) it is 0.92-0.97 versus 0.00-0.01. Keyed binning takes adaptive forgery from 100% to the false-positive floor at the primary operating point (bge, r=64). We also mark the boundary that guarantee does not cover: when the adversary copies the victim's own steps, shuffled splicing is neutralized (0.000 on Qwen2.5-3B) but chained replay remains at 0.76-0.98 across the five models, reported as open. Paraphrase robustness costs about half of the per-step watermark capacity. Code is available at https://anonymous.4open.science/r/SBW-Agent-Watermark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑