发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SkillClone黑盒攻击方法,可通过合法交互从隐藏的LLM智能体技能中重构其功能,表明仅文件保密无法确保功能保密。
AI 中文摘要
闭源智能体技能可能编码专有指令、脚本、常量和数据,提供商可将其能力作为服务提供,同时隐藏底层包。现有研究聚焦于直接泄露这些产物的提示注入攻击,对应的防御旨在防止此类泄露,但防止文件泄露并不能阻止用户恢复这些文件实现的功能。这引发了一个基本问题:在技能文件保持隐藏的情况下,用户能否通过常规使用重构技能的功能?本文研究行为技能重构(BSR),即攻击者利用合法任务请求和观测到的响应构建隐藏技能的功能克隆体。我们提出SkillClone,这是一种黑盒攻击方法,通过从目标技能的公开广告形成接口假设、发布结构化良性探测、合成可执行副本,以及通过与受害技能的差分验证迭代修复副本,来克隆目标技能。在涵盖规则、表格、流程和算法的30种技能上,SkillClone对多个目标的保留输入实现了精确或部分恢复,迭代重查询可弥合单轮重构遗漏的差距。由于SkillClone仅使用合法交互,以泄露为核心的防御覆盖范围有限,且技能描述越简略,保护效果越差。这些结果表明,仅文件保密无法确保功能保密,防御措施还必须限制常规使用带来的累积信息泄露。
英文摘要
Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection attacks that directly disclose these artifacts, and existing defenses accordingly aim to prevent such leakage. However, preventing file disclosure does not prevent users from recovering the functionality those files implement. This raises a fundamental question: can a user reconstruct a skill's functionality through ordinary use while its files remain hidden? We study behavioral skill reconstruction (BSR), in which an attacker uses valid task requests and observed responses to build a functional clone of a hidden skill. We introduce SkillClone, a black-box attack that clones a target skill by forming an interface hypothesis from its public advertisement, issuing structured benign probes, synthesizing an executable replica, and iteratively repairing it through differential validation against the victim skill. Across 30 skills spanning rules, tables, procedures, and algorithms, SkillClone achieves exact or partial recovery on held-out inputs for several targets. Iterative requerying closes gaps missed by single-round reconstruction. Because SkillClone uses only legitimate interactions, disclosure-focused defenses provide limited coverage, and less detailed skill descriptions offer limited protection. These results show that file secrecy alone does not ensure functional secrecy. Defenses must also limit cumulative information leakage from ordinary use.
Comments23 pages, 5 figures, 24 tables