arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34701cs.AIcs.LG

ResonAct:多智能体系统中用于运行时诊断与自愈的流式指标

ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems

Tarun Chintada, Neelamadhav Gantayat, Ishaan Romil, Renuka Sindhgatta, Soujanya Soni, Sameep Mehta

首次发表
浏览论文内容

中文总结 AI 辅助

ResonAct通过流式指标实现多智能体系统的运行时诊断与自愈,无需修改应用即可提升任务完成率最多10个百分点。

中文摘要 AI 辅助

多智能体系统(MAS)越来越多地被用于自动化涉及多个专业智能体、外部工具和长时间运行任务执行的企业工作流程。故障可能源于工具退化、上下文传播错误、协调中断或重复的智能体交互,这些都会阻碍任务完成。虽然现有的可观测性框架提供了追踪和日志,但诊断和修复大多在执行完成后进行,限制了运行时恢复的机会。我们提出了ResonAct,一个运行时自愈框架,通过流式操作指标实现对多智能体系统的持续监控、诊断和修复。ResonAct将执行追踪、智能体交互和工具调用摄入到流式分析层中,该层持续推导任务进度、上下文健康和工具可靠性指标。这些指标作为运行时控制信号,用于检测异常执行模式,并使用结构化故障模型定位根本原因。基于诊断出的故障,ResonAct动态选择修复策略并执行操作。该框架作为外部控制平面运行,无需修改应用智能体或编排逻辑即可进行干预。我们在企业工作流场景和AppWorld基准上评估了ResonAct。结果表明,基于流式指标的分析能够识别执行退化并定位故障。此外,策略驱动的修复将任务完成率提高了最多10.00个百分点,检测精度在70.59%到82.91%之间,召回率在63.09%到100%之间,恢复率在10.48%到46.67%之间,运行时开销在-0.25%到14.12%之间,覆盖了所有评估配置。

英文摘要

Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that prevent task completion. While existing observability frameworks provide traces and logs, diagnosis and remediation are largely performed after execution completes, limiting opportunities for recovery during runtime. We present ResonAct, a runtime self-healing framework that enables continuous monitoring, diagnosis, and remediation of multi-agent systems through streaming operational metrics. ResonAct ingests execution traces, agent interactions, and tool invocations into a streaming analytics layer that continuously derives task progress, context health, and tool reliability metrics. These metrics serve as runtime control signals for detecting anomalous execution patterns and localizing root causes using a structured failure model. Based on the diagnosed failure, ResonAct dynamically selects remediation policies and performs actions. The framework operates as an external control plane, enabling intervention without modifying application agents or orchestration logic. We evaluate ResonAct across enterprise workflow scenarios and AppWorld benchmarks. The results show that the streaming metric-based analysis identifies execution degradations and localizes faults. Furthermore, policy-driven remediation improves task completion rates by up to 10.00 percentage points, with detection precision ranging from 70.59% to 82.91%, recall from 63.09% to 100%, recovery rates from 10.48% to 46.67%, and runtime overhead ranging from $-0.25%$ to 14.12% across the evaluated configurations.

发表机构

  • IBM

机构由 AI 辅助整理,请以论文原文为准。

↑