arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

策略支撑的选择性再生:应对智能体间通信污染

Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication

Jinghan Xu, Longze Fan, Zeyuan Wang, Xinjin Li, Hankai Liu

arXiv 2609.26072首次发表:更新:

发表机构

Nankai University; China University of Petroleum; Sun Yat-sen University; Columnbia University(南开大学; 中国石油大学; 中山大学; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体通信中消息污染问题,提出ESC-CR框架,通过策略支撑的可执行承诺与洁净室恢复,在外部边界强制执行,选择性再生以保留证据信息并抑制未授权发布。

AI 中文摘要

智能体间通信对于多智能体语言模型系统至关重要,然而单条消息可能将任务关键信息与原始请求未授权的指令混合在一起。基于提示的防御将执行责任留给暴露于对抗性消息的模型,而不加区分的消息删除则会丢弃有用信息。我们提出可执行语义承诺与洁净室恢复(ESC-CR),一个策略支撑的框架,用于安全的智能体间代码生成与恢复。它将消息声明与授权分离,从可信任务、证据和策略构建可执行承诺,并在外部发布边界强制执行。一旦发生违规,ESC-CR将污染的责任消息和被拒绝的工件,从基于证据的任务信息重建洁净上下文,并在相同策略下重新生成。我们在通信关键型和标准代码生成基准、多种模型家族和通信拓扑,以及涵盖直接、混淆和验证器感知载荷的自适应攻击上评估ESC-CR。结果表明,污染上下文的重试经常无法消除未授权影响,而完全删除消息可能丢弃通信关键型任务所需的信息。在匹配的计算预算下,ESC-CR保留基于证据的声明,同时抑制未授权的发布,并且相同的设计可迁移到端到端的智能体轨迹。

英文摘要

Inter-agent communication is essential to multi-agent language-model systems, yet a single message may combine task-critical information with instructions not authorized by the original request. Prompt-based defenses leave enforcement to models exposed to adversarial messages, while indiscriminate message removal discards useful information. We introduce Executable Semantic Commitments with Clean-Room Recovery (ESC-CR), a policy-backed framework for secure inter-agent code generation and recovery. It separates message claims from authorization, constructs executable commitments from trusted tasks, evidence, and policy, and enforces them at an external release boundary. Upon a violation, ESC-CR taints the responsible message and rejected artifact, reconstructs a clean context from evidence-backed task information, and regenerates under the same policy. We evaluate ESC-CR across communication-essential and standard code-generation benchmarks, multiple model families and communication topologies, and adaptive attacks spanning direct, obfuscated, and verifier-aware payloads. Results show that polluted-context retry frequently fails to remove unauthorized influence, while complete message removal can discard information required by communication-essential tasks. ESC-CR preserves evidence-backed claims while suppressing unauthorized releases under matched computational budgets, and the same design transfers to end-to-end agent trajectories.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑