arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29015cs.AIcs.CLcs.DC

MeshHeal:去中心化LLM智能体网络中灰色故障的双时间尺度自愈

MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

Keru Chen, Sen Lin, Yingbin Liang, Nathaniel D. Bastian, Shaofeng Zou

首次发表
浏览论文内容

中文总结 AI 辅助

MeshHeal提出一种去中心化自愈框架,通过双时间尺度能力匹配评审检测并隔离灰色故障智能体,在保护当前任务的同时允许恢复后重新加入,显著提升退化阶段准确率并降低令牌开销。

中文摘要 AI 辅助

基于去中心化LLM的多智能体系统通过局部交互进行协调,但一个智能体可能在保持响应的同时,其任务解决质量持续下降。这种灰色故障需要在有足够证据改变未来路由之前保护当前任务,同时允许恢复的智能体重新加入。我们提出了MeshHeal,一个完全去中心化的自愈框架,它结合了跨两个时间尺度的能力匹配的同行评审。在快时间尺度上,一个自适应层级将不确定或低分的输出从重复的单评审者评估升级到委员会审议,并在必要时在使用前进行纠正。在慢时间尺度上,一个基于任务和能力条件的同行相对检测器聚合分数,以区分持续退化与普通输出变化,触发强制委员会审查,并最终将退化智能体从普通路由中排除;恢复探针为重新整合提供新证据。为了忠实评估路由,我们引入了基于模型的MAS评估,它将能力分配与执行模型绑定,因为仅基于提示的能力分配可能隐藏路由错误。在BBH、MATH和MMLUPRO上,MeshHeal在每任务使用51k总模型令牌的情况下实现了0.839的退化阶段准确率,而最强基线Symphony在每任务使用115k令牌的情况下实现了0.807的准确率。在交错退化和恢复下,MeshHeal隔离退化智能体,使其在恢复前不参与普通任务执行,并使其恢复正常路由。

英文摘要

Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive hierarchy escalates uncertain or low-scoring outputs from repeated single-reviewer evaluation to committee deliberation and, when needed, correction before use. At the slow timescale, a task- and ability-conditioned peer-relative detector aggregates scores to distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing; recovery probes provide fresh evidence for reintegration. To faithfully evaluate routing, we introduce Model-Backed MAS Evaluation, which ties ability assignments to execution models, since prompt-based ability assignments alone can leave routing errors hidden. Across BBH, MATH, and MMLU-Pro, MeshHeal achieves 0.839 degraded-phase accuracy using 51k total model tokens per task, versus the strongest baseline Symphony's 0.807 accuracy using 115k per task. Under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them excluded from ordinary task execution until recovery, and returns them to normal routing.

发表机构

  • Arizona State University(亚利桑那州立大学)
  • University of Houston(休斯顿大学)
  • The Ohio State University(俄亥俄州立大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑