arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07131cs.CRcs.AI

AgentLeak:超越技能窃取,将更强LLM智能体能力克隆到较弱智能体上

AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing

Xiaoting Lyu, Yuhong Wu, Yufei Han, Shichang Liu, Liang Zhang, Bin Wang, Bin Wang, Xiaobo Ma, Wei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

AgentLeak通过分析受害者与攻击者执行差异,识别能力关键行为并注入攻击者技能,在黑盒场景下将任务通过率提升超40%,恢复80%以上能力差距,揭示LLM智能体执行行为泄露程序性知识的机密性风险。

中文摘要 AI 辅助

大型语言模型(LLM)智能体通过将基础模型与显式技能以及通过执行获得的隐式程序性知识相结合,日益实现长期任务。由此产生的任务解决能力已成为有价值的专有资产,引发了一个新的安全问题:一个能力明显较弱的攻击者控制的智能体能否通过有限的黑盒交互获得较强专有智能体的能力?现有的技能窃取攻击能够恢复显式技能工件,但我们表明工件泄露并不一定转移能力:较弱的智能体可能拥有相同的技能,但仍会失败,因为它缺乏较强智能体隐式实现的程序性行为。我们的关键见解是,技能执行差距本身构成了一个泄露表面,其中缺失的行为通过成功的受害者执行与失败的攻击者执行之间的可观察差异暴露出来。基于此,我们提出了AgentLeak,一种黑盒能力克隆攻击,它从这些执行差异中识别能力关键行为,并将其纳入攻击者侧技能,同时保持攻击者的模型、框架和工具不变。在包含600个实例、多种智能体系统和多个骨干模型的20个任务场景中,与直接技能复用相比,AgentLeak将任务通过率提高了40%以上,并恢复了受害者与攻击者能力差距的80%以上。我们的发现揭示了LLM智能体中的机密性风险:仅保护显式工件是不够的,因为可观察的执行行为可能泄露重建低能力且攻击者控制的智能体中专有任务解决能力所需的程序性知识。

英文摘要

Large language model (LLM) agents increasingly achieve long-horizon tasks by combining foundation models with explicit skills and implicit procedural knowledge acquired through execution. The resulting task-solving capabilities have become valuable proprietary assets, raising a new security question: can a substantially weaker attacker-controlled agent acquire the capabilities of a stronger proprietary agent through limited black-box interaction? Existing skill-stealing attacks recover explicit skill artifacts, yet we show that artifact leakage does not necessarily transfer capability: a weaker agent may possess the same skills but still fail because it lacks procedural behaviors implicitly realized by the stronger agent. Our key insight is that the skill execution gap itself forms a leakage surface, where missing behaviors are exposed through observable differences between successful victim executions and failed attacker executions. Based on this, we present AgentLeak, a black-box capability-cloning attack that identifies capability-critical behaviors from these execution differences and incorporates them into attacker-side skills, while keeping the attacker's model, harness, and tools unchanged. Across 20 task scenarios comprising 600 instances, diverse agent systems, and multiple backbone models, AgentLeak improves task pass rates by over 40% compared with direct skill reuse and recovers more than 80% of the victim--attacker capability gap. Our findings reveal a confidentiality risk in LLM agents: protecting explicit artifacts alone is insufficient, as observable execution behavior can leak the procedural knowledge required to reconstruct proprietary task-solving capabilities in low-capability and attacker-controlled agents.

发表机构

  • Xi’an Jiaotong University(西安交通大学)
  • INRIA(法国国家信息与自动化研究所)
  • University of Warwick(华威大学)
  • Zhejiang Key Laboratory of Artificial Intelligence of Things (AIoT) Network and Data Security(浙江省人工智能物联网(AIoT)网络与数据安全重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑