arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信任门控能力控制:打破多智能体LLM系统中的信任-脆弱性悖论

Trust-Gated Capability Control: Breaking the Trust-Vulnerability Paradox in Multi-Agent LLM Systems

Mehedi Hasan Nipu, Chinmoy Mitra, Tarannum Ahmed Nowshin, Israt Moyeen Noumi

arXiv 2610.07000首次发表:更新:

发表机构

North South University; Rajshahi University of Engineering & Technology; BRAC University; Ahsanullah University of Science and Technology(北南大学; 拉杰沙希工程与技术大学; 布拉格大学; 阿赫桑努拉科学与技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体LLM系统中信任与脆弱性并存的悖论,提出可操作的五层信任堆栈,通过跨层协同、无遗憾权重估计和信任门控能力控制,实现可证明的权限安全与快速撤销。

AI 中文摘要

多智能体LLM系统的分层信任模型在很大程度上仍停留在概念层面:它们指出了信任的哪些维度重要,但没有说明各层如何组合、其重要性如何在运行时设定,或信任应如何支配智能体行为。这一缺口之所以重要,是因为更高的智能体间信任在提高任务成功率的同时,也扩大了被利用的风险敞口,这种张力被形式化为“信任-脆弱性悖论”。我们通过三项贡献使五层信任堆栈变得可操作。首先,一个跨层协同算子将前置层的缺陷传播到依赖层,馈入一个广义均值复合信任,该信任在极限情况下恢复最弱环节规则,并具有可证明的界。其次,各层的重要性权重通过一个无遗憾在线估计器基于观测到的失败进行设定,该估计器追踪当前对危害负最大责任的层。第三,信任门控能力控制仅在复合信任及相关前置层通过能力特定阈值时,发放短期、可撤销的能力授权。我们证明该机制打破了这一悖论:一种通过夸大行为信任同时削弱前置层的隐蔽入侵无法升级权限,并给出了一个闭式、可证明保守的信任不动点,以及有界延迟的撤销保证。一项带有潜伏对手的数值研究证实了这些预测:复合信任界在20,000次随机抽取中成立,解析不动点与模拟的差异在0.007以内,且受损的前置层会在几次交互内触发自动撤销。

英文摘要

Layered trust models for multi-agent LLM systems remain largely conceptual: they name which dimensions of trust matter but not how layers combine, how their importance is set at runtime, or how trust should govern agent actions. This gap matters because higher inter-agent trust raises task success while also enlarging exposure to exploitation, a tension formalized as the Trust-Vulnerability Paradox. We make a five-layer trust stack operational through three contributions. First, a cross-layer synergy operator propagates prerequisite-layer deficits into de- pendent layers, feeding a generalized-mean composite trust that recovers the weakest-link rule as a limiting case, with provable bounds. Second, per-layer importance weights are grounded in observed failures via a no-regret online estimator that tracks which layer is currently most responsible for harm. Third, Trust- Gated Capability Control issues short-lived, revocable capability grants only when composite trust and the relevant prerequisite layers clear capability-specific thresholds. We prove this mechanism breaks the paradox: a stealthy compromise that inflates behavioral trust while degrading a prerequisite layer cannot escalate privilege, and give a closed-form, provably conservative trust fixed point with a bounded-latency revocation guarantee. A numerical study with a sleeper adversary confirms the predictions: composite-trust bounds hold across 20,000 random draws, the analytic fixed point matches simulation within 0.007, and a compromised prerequisite layer triggers automatic revocation within a few interactions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑