LoRo-Mark:可证明无损且鲁棒的智能体水印
LoRo-Mark:Provably Lossless and Robust Agent Watermarking
浏览论文内容
中文总结 AI 辅助
针对LLM智能体重打包侵权问题,提出LoRo-Mark水印机制,通过密码学认证的法证分支实现可证明无损,并冗余分布所有权信息以增强鲁棒性,实验验证零性能损失与可靠验证。
中文摘要 AI 辅助
随着LLM智能体日益作为商业服务部署,保护专有的编排逻辑和工具使用策略变得重要。我们考虑智能体重打包问题:对手通过API将受保护的智能体集成到自己的应用程序中,并以自己的身份呈现。它可能修改部分执行以掩盖来源。所有者通常只能对重打包服务进行黑盒访问,因此黑盒所有权验证至关重要。智能体水印将所有权证据嵌入智能体行为中,以便后续验证。有效的水印应满足两个要求:无损性,即保持原始功能;鲁棒性,即在执行部分修改后仍能恢复所有权证据。现有方法通常将信号嵌入行为选择或执行轨迹中,干预正常决策,并且对行为修改的鲁棒性有限。我们提出LoRo-Mark,一种可证明无损且鲁棒的智能体水印机制。对于无损性,它将水印隔离到密码学认证的法证分支中,该分支在正常执行期间保持不活跃,仅由所有者授权的请求激活。通过将未授权分支激活减少到标准MAC安全性,LoRo-Mark正式保证性能保持。对于鲁棒性,它在法证行为序列中冗余分布所有权信息,使得在部分行为替换和序列截断下能够可靠恢复。在多个LLM智能体上的实验显示,正常任务零退化,并且在序列修改下可靠的所有权验证。
英文摘要
As LLM agents are increasingly deployed as commercial services, protecting proprietary orchestration logic and tool-use policies is important. We consider agent repackaging: an adversary integrates a protected agent into its own application via API and presents it under its own identity. It may modify parts of execution to obscure the source. The owner typically has only black-box access to the repackaged service, so black-box ownership verification is essential. Agent watermarking embeds ownership evidence into agent behavior for later verification. An effective watermark should satisfy two requirements: losslessness, preserving original functionality, and robustness, keeping ownership evidence recoverable after partial modification of execution. Existing methods often embed signals into behavior selection or execution trajectories, intervene in normal decisions, and offer limited robustness to behavior modification. We propose LoRo-Mark, a provably lossless and robust agent watermarking mechanism. For losslessness, it isolates watermarking into a cryptographically authenticated forensic branch that remains inactive during normal execution and is activated only by owner-authorized requests. By reducing unauthorized branch activation to standard MAC security, LoRo-Mark formally guarantees performance preservation. For robustness, it redundantly distributes ownership information across forensic behavior sequences, enabling reliable recovery under partial behavior substitution and sequence truncation. Experiments across multiple LLM agents show zero degradation on normal tasks and reliable ownership verification under sequence modifications.
发表机构
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。