AI 中文总结
研究多智能体系统恶意传播问题,提出SafeFlow框架,将恶意跨智能体传播形式化为语义信息流问题,通过附加污点、传播及工作流级验证重建风险上下文,经测试降低攻击成功率,保持良性任务完成率,揭示系统缺乏跨委托边界保留风险语义机制。
AI 中文摘要
多智能体系统通过任务分解和角色专业化提高能力,但这些机制引入了安全盲点:有害目标可分解为局部合理的子任务,使恶意意图逃避单个智能体检测。这是一个日益严峻的社会影响挑战。我们认为应将此失败模式理解为语义信息流问题而非单轮提示分类任务。为此,我们提出SafeFlow,一个将恶意跨智能体传播形式化为语义信息流问题的多智能体系统防御框架。SafeFlow给根请求附加结构化语义污点,通过动态协作图传播,在执行不可逆操作前进行工作流级验证以重建全局风险上下文。在四个基准测试上评估,SafeFlow与无防御基线和外部防御相比降低了攻击成功率,同时保持高良性任务完成率和高配对安全-危害成功率。我们的发现表明多智能体系统仍缺乏跨委托边界保留风险语义的机制,SafeFlow在危害发生前使风险在整个工作流中可见。
英文摘要
Multi-agent systems improve capability through task decomposition and role specialization, but these same mechanisms introduce an important safety blind spot: a harmful objective can be fragmented into locally plausible subtasks, allowing malicious intent to evade detection by any single agent. This is a growing social-impact challenge: systems handling sensitive information or consequential tools can turn routine delegation into unauthorized disclosure or unsafe action. We argue that this failure mode is better understood as a semantic information-flow problem than as a single-turn prompt classification task. To address this, we propose SafeFlow, a defense framework for multi-agent systems that formalizes malicious cross-agent propagation as a semantic information-flow problem. SafeFlow attaches structured semantic taints to root requests, propagates them through a dynamic collaboration graph, and performs workflow-level validation to reconstruct the global risk context before irreversible actions are committed. Evaluated on four benchmarks spanning prompt injection, jailbreak-based unsafe tool use, risky code execution, and harmful web-agent behavior, SafeFlow reduces attack success rates compared to undefended baselines and external defenses while retaining high benign task completion and a high paired safe--harm success rate. Our findings show that multi-agent systems still lack mechanisms for preserving risk semantics across delegation boundaries. This gap can turn routine delegation into privacy harms or unsafe actions that affect people and organizations. SafeFlow keeps this risk visible throughout the workflow, before it results in harm.