arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16026cs.CR

SkillWatermark:通过良性提示实现渐进式隐私推理的嵌入式技能水印

SkillWatermark: An Embedded Skill Watermark of Progressive Privacy Inference via Benign Prompts

Yu Li, Liqi Zhuang, Dong Wei, Jiwen Luo, Hang Zhang, Meng Zhang, Xiaona Li, Weiqing Huang

AI总结:

本研究提出SkillWatermark,通过良性提示插入技能水印在LLM智能体多轮对话流量中编码私密信息,经实验验证水印效果良好且能通过安全审计,揭示了流量模式可被利用的新型攻击面。

AI中文摘要:

大型语言模型(LLM)智能体的技能已广泛部署于各类应用领域,但研究发现这些技能在执行过程中会产生特定流量模式。本文设计了一种生成特定流量模式的流程,通过插入精心设计的技能描述(称为技能水印),使被动网络攻击者能建立隐蔽信道,在多轮对话的可观测流量中编码私密信息。具体而言,我们将称为水印的提示约束术语插入原始技能描述,并嵌入多轮对话中,用户原始提示中的关键信息会被这些水印触发,在流量中产生清晰可观测的编码。攻击者只需解码流量模式即可恢复编码信息。特别地,我们的修改是良性的,不会直接泄露任何私密数据,也不会执行任何恶意指令。大量实验表明,我们的水印能产生高度一致且可区分的流量模式,且经修改的技能可通过现有基于LLM的安全审计工具。本研究强调,生成特定流量模式可被利用为新型攻击面,并为未来的安全强化提供关键见解。

英文摘要:

Skills for large language model (LLM) agents have been widely deployed across diverse application domains. However, we observe that these skills generate specific traffic patterns during execution. In this paper, we design a pipeline that generates specific traffic patterns by inserting carefully designed skill descriptions, which we term skill watermarks, so that a passive network attacker can establish a covert channel to encode private information within observable traffic across multiple conversation turns. Specifically, we insert prompt constraint terms, referred to as watermarks, into the original skill descriptions and embed them within multi-turn conversations. The key information in the user's original prompt is thereby triggered by these watermarks, producing clearly observable encodings in the traffic. The adversary need only decode the traffic patterns to recover the encoded information. In particular, our modifications are benign in the sense that they do not directly exfiltrate any private data and do not execute any malicious instructions. Extensive experiments demonstrate that our watermarks produce highly consistent and distinguishable traffic patterns, and that the transformed skills pass existing LLM-based security auditing tools. This study highlights that generating specific traffic patterns can be exploited as a novel attack surface and offers critical insights for future security hardening.

↑