发表机构
UC Berkeley; National Yang Ming Chiao Tung University(加州大学伯克利分校; 国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Daydreaming攻击,通过黑盒任务交互窃取智能体技能,在7种技能和4种模型上恢复了原始技能86.8%的能力,性能优于SigLeak,证明仅隐藏技能文件无法阻止功能重建。
AI 中文摘要
智能体技能捆绑了指令、参考数据和可执行辅助工具,使通用智能体能够执行专门任务。托管提供商可在出售任务结果访问权限的同时将这些文件保密,这使得技能本身成为有价值的目标。现有的泄露防御措施可以阻止要求获取技能或复制其文本的请求,但无法阻止客户提交服务旨在完成的普通任务。我们提出了Daydreaming,这是一种仅执行的攻击,通过黑盒任务交互窃取多文件技能。受害者从未被要求泄露技能或对重建结果进行评分。相反,Daydreaming自适应地创建精心设计的任务,其结果可区分可能的隐藏行为。它测试单个行为,使用攻击者控制的影子智能体选择设计,并使用存储的受害者结果和本地执行检查完成每个文件。我们将三种嵌套的访问威胁级别形式化为差分、跟踪和输出,并专注于输出级别,其中攻击者仅看到最终响应和返回的文件。在7种技能和4种受害者模型上,Daydreaming在输出级别恢复了原始技能86.8%的能力,比SigLeak的性能高出近4倍。即使启用了泄露防御,它每个技能平均仅使用32次受害者调用即可生成可安装的技能。这些结果表明,隐藏技能文件和过滤直接泄露本身并不能阻止通过正常使用进行的功能重建。
英文摘要
Agent skills bundle instructions, reference data, and executable helpers that let a general agent perform specialized tasks. Hosted providers can keep these files secret while selling access to task results, making the skill itself a valuable target. Existing disclosure defenses can block requests that ask for the skill or reproduce its text, but they cannot block customers from submitting the ordinary tasks the service is built to complete. We present Daydreaming, an execution-only attack that steals a multi-file skill through black-box task interactions. The victim is never asked to reveal the skill or grade a reconstruction. Instead, Daydreaming adaptively creates crafted tasks whose results distinguish possible hidden behaviors. It tests individual behaviors, uses attacker-controlled shadow agents to choose a design, and completes each file using stored victim results and local execution checks. We formalize three nested threat levels of access as Differential, Trace, and Output, and focus on Output, where the attacker sees only the final response and returned files. Across 7 skills and 4 victim models, Daydreaming recovers 86.8% of the original skill's capability at Output, outperforming SigLeak by almost 4x. It produces installable skills using a median of 32 victim calls per skill even with disclosure defenses enabled. These results show that hiding skill files and filtering direct disclosure do not, by themselves, prevent functional reconstruction through normal use.