arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体调控蒸馏:自主多智能体系统中的推理时调控提取与利用

Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems

Yu Cui, Wuli Yang, Yirui Shi, Junhao Xia, Hui Jiang, Lei Gao, Chenfu Bao

arXiv 2607.28147首次发表:更新:

AI 中文总结

本文提出AHD框架,将AMAS推理时调控提取形式化为新安全问题,实验证实其可提取调控致IP泄露,还提出对应欺骗防御,揭示AMAS新安全威胁。

AI 中文摘要

基于大语言模型(LLM)构建的自主多智能体系统(AMAS,如Hermes)日益依赖推理时调控(inference-time harnesses)来协调推理与行动。构建此类调控需要大量工程投入与计算资源,因其需在组合搜索空间中迭代优化,且会随底层LLM共同演化,故推理时调控构成有价值的知识产权(IP)。尽管已有研究探讨了静态多智能体系统(具预配置架构)中的IP泄露风险,但AMAS中是否存在类似风险仍不明确——AMAS的调控行为在推理过程中动态生成。为填补该空白,本文提出Agent Harness Distillation(AHD)框架,用于研究AMAS中推理时调控提取引发的安全风险。我们将调控提取形式化为新安全问题,并开发评估框架以量化此类风险。AHD通过黑盒交互从目标智能体提取推理时调控能力,包含两个阶段:预蒸馏阶段,AHD从目标智能体的响应中推断推理时调控行为,构建初始调控;后蒸馏阶段,AHD迭代优化初始调控,使其与目标智能体的行为模式对齐。在真实AMAS(跨多个骨干LLM)上的实验表明,AHD有效,且存在显著IP泄露风险。我们进一步提出基于欺骗的防御措施,可在保留受保护智能体效用的同时降低调控提取有效性。研究结果揭示了AMAS此前未被充分探索的安全威胁。

英文摘要

Autonomous multi-agent systems (AMAS) built on large language models (LLMs), such as Hermes, increasingly rely on inference-time harnesses to coordinate reasoning and action. Constructing these harnesses requires substantial engineering effort and computational resources, as they are iteratively optimized over a combinatorial search space while co-evolving with the underlying LLM. Inference-time harnesses therefore constitute valuable intellectual property (IP). Although prior work has investigated IP leakage in static multi-agent systems with pre-configured architectures, it remains unclear whether similar risks arise in AMAS, where harness behavior emerges dynamically during inference. To address this gap, we introduce Agent Harness Distillation (AHD), a framework for studying the security risks arising from inference-time harness extraction in AMAS. We formalize harness extraction as a new security problem and develop an evaluation framework for quantifying such risks. AHD extracts inference-time harness capabilities from a target agent through black-box interactions and consists of two stages. In the pre-distillation stage, AHD infers inference-time harness behaviors from the responses of the target agent and constructs an initial harness. In the post-distillation stage, AHD iteratively refines the initial harness to align with the behavioral patterns of the target agent. Experiments on real-world AMAS across multiple backbone LLMs demonstrate the effectiveness of AHD and reveal substantial IP leakage risks. We further propose a deception-based defense that reduces harness extraction effectiveness while preserving the utility of the protected agent. Our findings uncover a previously underexplored security threat to AMAS.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑