发表机构
Nankai University; China University of Petroleum; Sun Yat-sen University; Columnbia University(南开大学; 中国石油大学; 中山大学; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多智能体LLM系统的隐私泄露问题,提出范围受限语义解密协议MNC,通过绑定披露范围实现授权信息传递,阻止未授权操作,为私有LLM智能体系统提供实用通信边界。
AI 中文摘要
多智能体大语言模型(LLM)系统即便公开输出看似无害,也可能通过内部消息、工具参数、日志及持久内存暴露受保护状态。现有的隐私提示、编辑方法和源代码级访问控制仅限制表面内容或数据访问,未明确合法知情智能体应披露什么,以及该披露信息如何在下游复用。我们提出最小必要通信(Minimum-Necessary Communication, MNC),这是一种类型化语义解密协议,可从应用开发者编写的候选集合中选择满足任务需求的披露信息,并将其绑定到明确的接收方、目的、转发权限、生命周期、日志记录及内存范围。参考监视器会在后续操作中强制执行这些范围,而感知历史的扩展机制则考虑了多次披露过程中累积的推理风险。受控语义连接、内存探测及纵向实验表明,传统防御机制虽能保留协议级效用,但会暴露大量额外推理信号。在相同接收文本下,MNC在保留授权传递的同时,可阻止纯文本语义解密器允许的未授权转发、日志记录、持久存储及过期后检索。双主干MAGPIE执行进一步显示,中介披露会通过后续规划、工具使用、协调及内存检索传播。这些结果支持范围受限语义解密作为私有LLM智能体系统的实用通信边界。
英文摘要
Multi-agent large language model (LLM) systems can expose protected state through internal messages, tool arguments, logs, and persistent memory even when their public outputs appear innocuous. Existing privacy prompts, redaction methods, and source-level access controls restrict surface content or data access, but do not specify what a legitimately informed agent should disclose or how that disclosure may be reused downstream. We introduce Minimum-Necessary Communication (MNC), a typed semantic-declassification protocol that selects a task-sufficient disclosure from an application-authored candidate family and binds it to explicit recipient, purpose, forwarding, lifetime, logging, and memory scopes. A reference monitor enforces these scopes across subsequent operations, while a history-aware extension accounts for inference risk accumulated over repeated disclosures. Controlled semantic-join, memory, probing, and longitudinal experiments show that conventional defenses can preserve protocol-level utility while exposing substantial additional inference signal. Under identical receipt text, MNC preserves authorized delivery while blocking unauthorized forwarding, logging, durable storage, and retrieval after expiration that a text-only semantic declassifier permits. Two-backbone MAGPIE executions further show that mediated disclosures propagate through subsequent planning, tool use, coordination, and memory retrieval. These results support scope-bound semantic declassification as a practical communication boundary for private LLM-agent systems.