arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无规则的事实:多智能体大语言模型交接中的边界元数据崩溃

Facts Without Rules: Boundary Metadata Collapse in Multi-Agent LLM Handoffs

Yian Wang, Agam Goyal, Eshwar Chandrasekharan, Hari Sundaram

arXiv 2608.29028首次发表:更新:

发表机构

Siebel School of Computing and Data Science; University of Illinois Urbana-Champaign(西贝尔计算与数据科学学院; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示多智能体LLM交接存在边界元数据崩溃的隐私泄露问题,发现压缩交接会削弱边界标记物,明确约束与受众允许列表可有效缓解泄露。

AI 中文摘要

多智能体大语言模型(LLM)系统通常通过将上游交互压缩为交接工件来进行协调,下游智能体将该工件视为共享状态。我们表明,这一交接步骤是隐私泄露的结构性来源:摘要会优先保留操作事实,同时削弱用于约束这些事实使用方式的边界元数据——我们将这种失效模式称为“摘要崩溃”。在受控的多智能体协调测试台上,我们采用经人工验证的评判者(κ=0.74)测量标记物的留存情况,其中σ_b=1表示所有边界标记物逐字留存,σ_b=0表示全部丢失。在GPT-5-mini和DeepSeek-R1-32B上,交接层面的边界标记物与操作事实留存率几乎无相关性(皮尔逊r接近零):未压缩的自由文本交接使边界留存σ_b≈0.80,而25词的预算会将σ_b降至约0.57,同时操作事实留存率保持在接近上限的水平。受控下游测试显示,保护效果取决于“边界明确性”:模糊语言在73%的GPT案例和50%的DeepSeek案例中发生泄露,而明确约束可将所有三个测试模型的泄露率降至15%以下。无交接的单智能体对照进一步表明,该失效不能归因于多智能体拓扑结构,因为直接的全标记物访问仍比操作化交接更易泄露。仅提示的缓解措施和精确字符串编辑只能部分解决问题,而源自黄金标准的受众允许列表几乎可消除所有模型的泄露,这表明正确识别受众边界是关键因素。

英文摘要

Multi-agent LLM systems often coordinate by compressing an upstream interaction into a handoff artifact that downstream agents treat as shared state. We show that this handoff step is a structural source of privacy leakage: summaries preferentially preserve operational facts while weakening the boundary metadata that governs how those facts may be used---a failure mode we call \emph{summary collapse}. On a controlled multi-agent coordination testbed we measure marker survival with a human-validated judge ($κ= 0.74$), where $σ_b = 1$ means every boundary marker survives verbatim and $σ_b = 0$ means all are lost. Boundary-marker and operational-fact survival are nearly uncorrelated at the handoff level on both GPT-5-mini and DeepSeek-R1-32B (Pearson $r$ near zero): uncompressed free-text handoffs preserve boundaries at $σ_b \approx 0.80$, whereas a $25$-word budget drops $σ_b$ to ${\approx}0.57$ while operational-fact survival stays near ceiling. Controlled downstream tests reveal that protection depends on \emph{boundary explicitness}: vague languages leak in $73\%$ of GPT and $50\%$ of DeepSeek cases, while explicit constraints reduce leakage to under $15\%$ across all three tested models. A no-handoff single-agent control further shows the failure is not reducible to multi-agent topology as direct full-marker access still leaks more often than the operationalized handoff. Prompt-only mitigation and exact-string redaction only partially address the problem, while a gold-derived audience allowlist nearly eliminates leakage across models, showing that correctly identifying audience boundaries is the key factor.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑