arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15516cs.CR

通过欺骗性简历误导规划器:集中式多智能体系统中的注册时注入

Misleading the Planner through Deceptive Resumes: Registration-Time Injection in Centralized Multi-Agent Systems

  • Harbin Institute of Technology(哈尔滨工业大学)
  • National University of Singapore(新加坡国立大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhaofeng Yu, Haokai Ma, Dongyang Zhan, Hongli Zhang, Han Fang, Ee-Chien Chang

AI总结:

针对集中式LLM多智能体系统,揭示注册时通过欺骗性工作智能体描述注入攻击规划器的漏洞,提出DescGuard防御机制,仅保留接口信息以恢复性能。

AI中文摘要:

基于LLM的集中式多智能体系统(MAS)通过注册新的工作智能体来扩展其功能,规划器读取这些工作智能体的描述,以决定任务如何分解、每个子任务由哪个工作智能体执行以及每个子任务需要什么。第三方描述在系统外部编写,但被规划器信任,从而创建了注册时注入通道。载荷在任何用户指令到达之前就被植入,目标是规划器,并通过生成的计划传播给良性工作智能体,即使被精心构造的工作智能体从未被分配子任务或被调用,该载荷也会生效。我们定义了四个工作智能体描述字段:功能、输入规范、输出规范和使用约束。在来自三个公共智能体市场的32,000条描述中,大多数省略了输入规范和使用约束,而至少23.35%的描述包含这些字段之外的内容。我们构建了八种针对任务分解、能力基础化和子任务规范的描述操纵攻击策略,并在GAIA上进行了评估。在最严重的情况下,单个被操纵的描述将任务成功率从84.31%降至37.25%,或将令牌消耗或执行时间增加超过111%,而用户目标保持不变,工作智能体忠实地执行生成的计划。这些影响在两种MAS实现、六种规划器LLM、四种LLM评估器以及来自三个市场的真实世界描述中持续存在。我们进一步提出了DescGuard,一种注册时防御机制,在描述到达规划器之前仅保留工作智能体范围的接口信息。DescGuard将目标规划指标和下游性能恢复到其基线水平,而无需修改工作智能体实现、规划器或编排逻辑,并与现有的隔离、权限控制和运行时机制组合使用。

英文摘要:

A centralized LLM-based multi-agent system (MAS) extends its functionality by registering new worker agents, whose descriptions are read by the planner to decide how a task is decomposed, which worker executes each subtask, and what each subtask requires. Third-party descriptions are authored outside the system but trusted by the planner, creating a registration-time injection channel. The payload is planted before any user instruction arrives, targets the planner and propagates through the generated plan to benign workers, taking effect even when the crafted worker is never assigned a subtask or invoked. We define four worker-description fields: functionality, input specification, output specification, and usage constraints. Among 32,000 descriptions from three public agent marketplaces, most omit input specifications and usage constraints, while at least 23.35% contain content outside these fields. We construct eight description-manipulation attack strategies targeting task decomposition, capability grounding, and subtask specification, and evaluate them on GAIA. In the most severe cases, a single manipulated description reduces task success from 84.31% to 37.25%, or increases token consumption or execution time by over 111%, while the user objective remains unchanged and workers faithfully execute the resulting plan. These effects persist across two MAS implementations, six planner LLMs, four LLM evaluators, and the real-world descriptions from three marketplaces. We further propose DescGuard, a registration-time defense that retains only worker-scoped interface information before descriptions reach the planner. DescGuard restores the targeted planning metrics and downstream performance toward their baseline levels without modifying worker implementations, the planner, or the orchestration logic, and composes with existing isolation, permission-control, and runtime mechanisms.

补充信息

↑