发表机构
Xidian University; Zhejiang University; Sun Yat-sen University; Shanghai Jiao Tong University; HKUST(西安电子科技大学; 浙江大学; 中山大学; 上海交通大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体轨迹被非法蒸馏的问题,提出AuxMark行为水印框架,通过插入辅助动作并利用证据卡进行黑盒审计,实现高精度检测与归因,且保持任务效用。
AI 中文摘要
大型语言模型智能体可以通过多步交互和工具使用获得复杂能力,但其轨迹也可能被非法收集以蒸馏出学生智能体。然而,现有的水印方法要么不适合智能体环境的结构化和交互性,要么在跨任务和模型架构上缺乏可靠的效力。我们提出了AuxMark,一种用于追踪未经授权的智能体蒸馏的行为水印框架。AuxMark动态地将安全、非必要的辅助动作插入教师轨迹中,并将相关上下文存储为私有证据卡。为了审计可疑的学生模型,AuxMark从这些卡构建成对的真实和虚假探针,并应用卡级符号检验。这种黑盒协议支持模型级检测和轨迹级归因。在三个智能体基准、两个教师智能体和四个学生架构上,AuxMark检测出所有24个蒸馏模型,并在48个干净模型上实现了零误报。它还保持了任务效用,并对数据泛滥、释义、截断和自适应清理攻击保持有效。我们的代码将在该URL发布。
英文摘要
Large language model agents can acquire complex capabilities through multi-step interaction and tool use, but their trajectories can also be illegally collected to dis- till student agents. However, existing watermarking methods either do not fit the structured and interactive nature of agent environments or lack reliable effective- ness across tasks and model architectures. We introduce AuxMark, a behavioral watermarking framework for tracing unauthorized agent distillation. AuxMark dynamically inserts safe, non-essential auxiliary action into teacher trajectories, and stores the associated contexts as private evidence cards. To audit a suspicious student model, AuxMark constructs paired real and fake probes from these cards and applies a card-level sign test. This black-box protocol supports both model- level detection and trace-level attribution. Across three agent benchmarks, two teacher agents, and four student architectures, AuxMark detects all 24 distilled models with zero false positives on 48 clean models. It also preserves task utility and remains effective against data flooding, paraphrasing, truncation, and adaptive cleaning attacks. Our code will be released at this URL.
Comments24 pages