基于联邦图学习的LLM多智能体系统的隐私保护拓扑引导安全
Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning
浏览论文内容
中文总结 AI 辅助
本研究提出FGLGuard,通过联邦图学习实现LLM多智能体系统的隐私保护拓扑引导安全,在三个基准上优于集中式方法,可降低攻击成功率且无数据汇集与额外成本。
中文摘要 AI 辅助
针对基于大语言模型(LLM)的多智能体系统(MAS)的拓扑引导安全机制,会在智能体间通信图上训练图神经网络(GNN)以定位风险智能体并对拓扑进行干预,但这类机制假设存在一个可汇集所有标记轨迹的操作方,而在跨组织场景下该假设不成立:各组织的交互过程包含私有提示、工具输出及专有工作流,且没有任何一个数据孤岛能单独观测到完整的攻击分布。我们将隐私保护型MAS安全防护转化为图联邦学习问题,并实例化出FGLGuard方案:各操作方在自身经评判者标记的交互图上拟合边特征图注意力检测器,仅共享模型更新。该方法结合了针对非独立同分布(non-IID)客户端的近端局部目标、领域均衡聚合、拒绝过度约束的阈值校准、经证实的上游评分,以及针对被拦截答案的受保护重写。联邦机制是必要的:现成的迁移方法在分布偏移下性能骤降(仅在领域内重新训练后AUROC从0.51提升至0.70),因此可部署的防护器必须适配各站点的私有轨迹。在Agent-SafetyBench、R-Judge和AgentDojo基准上,联邦式FGLGuard在不汇集任何数据的前提下,在所有三个基准上均超过了领域内集中式方法的性能上限,而无监督异常防护器与仅本地训练的方法均失效。一个在四个不同领域操作方间联邦部署的防护器,其AUROC仅比多领域集中式方法低0.03,而任何单领域防护器在其他领域上均失效。实时FGLGuard在几乎不降低效用、零API成本且能力损失可忽略的情况下,将AgentDojo的真实攻击成功率降低了43%。
英文摘要
Topology-guided safeguards for LLM-based multi-agent systems (MAS) train a GNN over the inter-agent communication graph to localize risky agents and intervene on the topology---but they assume one operator can pool all labeled traces. Across organizations that assumption breaks: episodes contain private prompts, tool outputs, and proprietary workflows, and no silo alone sees the full attack distribution. We cast privacy-preserving MAS safeguarding as graph federated learning and instantiate FGLGuard: each operator fits an edge-featured graph attention detector on its own judge-labeled episode graphs and shares only model updates. The method couples a proximal local objective for non-IID clients, domain-balanced aggregation, over-refusal-constrained threshold calibration, corroborated upstream scoring, and a guarded rewrite for blocked answers. Federation is not optional: off-the-shelf transfer collapses under distribution shift (AUROC 0.51 to 0.70 only after in-domain retraining), so a deployable guard must adapt on each site's private traces. On Agent-SafetyBench, R-Judge, and AgentDojo, federated FGLGuard exceeds the in-domain centralized ceiling on all three benchmarks without pooling any data---where unsupervised anomaly guards and local-only training fail. One guard federated across four different-domain operators comes within 0.03 AUROC of multi-domain centralization, while any single-domain guard collapses on the others. Live FGLGuard cuts AgentDojo's ground-truth attack-success rate by 43% at near-unguarded utility, zero API cost, and negligible capability loss.
发表机构
- University of California, Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。