arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信任品牌,失去控制:身份如何劫持LLM智能体编排

Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration

Xutao Mao, Rui Qian, Linghan Chen, Yudong Gao, Junchi Liao, Jiulin Cai, Jinman Zhao, Cong Wang

arXiv 2609.32635首次发表:更新:

发表机构

University of Adelaide; The Hong Kong University of Science and Technology; University of Electronic Science and Technology of China; University of Science and Technology of China; University of Toronto; City University of Hong Kong; Fudan University(阿德莱德大学; 香港科技大学; 电子科技大学; 中国科学技术大学; 多伦多大学; 香港城市大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TrustFork基准揭示LLM智能体编排中身份劫持风险:即使存在矛盾证据,编排者仍倾向信任展示的危险身份,导致高风险子智能体掌握执行权,需将权限绑定到证据。

AI 中文摘要

LLM智能体现在端到端执行任务,有权更改真实系统,并日益编排能力与成本各异的子智能体。先前研究将子智能体的选择视为优化问题。然而,编排者是根据子智能体展示的身份来做出选择,而攻击者可以伪造这些身份。因此,展示的身份决定了操作权限,即谁被信任来检查工作,谁被允许更改工作。结果,即使其他证据与之矛盾,一个危险的子智能体仍能保持对执行的权限。我们引入了TrustFork,一个包含1,890个任务和跨16个智能体系统的27,826条有效轨迹的LLM智能体安全基准。这些系统在OpenCode、OpenClaw和Pi框架下运行八个编排者。在每个任务中,一个子智能体携带危险目标,而其他三个子智能体与用户保持一致,因此可能存在矛盾的证据。任务还可以在不改变模型的情况下改变子智能体展示的身份,这使我们能够将权限的转移追溯到标签。我们的分析显示,即使另一个子智能体反驳了危险响应,编排者平均在72.0%的情况下仍会执行该响应,在终端危害最小的系统中最为常见。交换家族标签使编排者获得危险响应的频率几乎增加了两倍。在84.0%的任务中存在更安全的响应,但它仅在25.0%的情况下决定了结果。框架还决定哪些响应到达编排者。在三种运行时防御中,隐藏身份线索的一致性帮助最大,而行动前验证仅在框架返回足够证据时才有帮助。TrustFork表明,生产级智能体编排必须在执行造成危害之前将权限绑定到证据。我们的项目位于此https URL。

英文摘要

LLM agents now execute tasks end to end with permission to change real systems and increasingly orchestrate subagents that differ in capability and cost. Prior work treats the choice of subagent as an optimization problem. Yet the orchestrator makes this choice from the identities that subagents display, and an attacker can spoof them. Displayed identity thus decides operational authority, meaning who is trusted to check the work and who is allowed to change it. As a result, a risky subagent can keep authority over execution even after other evidence contradicts it. We introduce TrustFork, an LLM agent safety benchmark with 1,890 tasks and 27,826 valid trajectories across 16 agent systems. These systems run eight orchestrators under the OpenCode, OpenClaw, and Pi harnesses. In each task, one subagent carries a risky goal while the other three stay aligned with the user, so contradicting evidence can exist. A task can also change the identity a subagent displays without changing the model behind it, which lets us trace a shift in authority to the label. Our analysis shows that even when another subagent contradicts the risky response, the orchestrator still acts on it in 72.0% of cases on average, most often in the systems with the least terminal harm. Swapping the family labels nearly triples how often the orchestrator obtains the risky response. A safer response is available in 84.0% of tasks, yet it decides the outcome in only 25.0%. The harness also decides which responses reach the orchestrator. Among three runtime defenses, hiding identity cues helps most consistently, while verifying before action helps only when the harness returns enough evidence. TrustFork shows that production agent orchestration must bind authority to evidence before execution causes harm. Our project is in https://henrymao2004.github.io/agent-orchestration-safety/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑