arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

软件工程中基于大语言模型的多智能体系统的可观测性与故障注入

Observability and Fault Injection for LLM-Based Multi-Agent Systems in Software Engineering

Zahra Seyedghorban, Egor Klimov, Arie van Deursen, Annibale Panichella, Burcu Kulahcioglu Ozkan

arXiv 2608.24271首次发表:更新:

AI 中文总结

针对软件工程中基于LLM的多智能体系统难检查调试评估的问题,提出llmmas-otel工具,结合OpenTelemetry分布式追踪与故障注入,实现可复现的基线与故障执行对比,已在演示及真实系统上验证。

AI 中文摘要

基于大语言模型(LLM)的多智能体系统正被越来越多地探索用于软件工程任务,但在受控故障下仍难以检查、调试和评估。我们提出llmmas-otel,这是一种轻量级、与框架无关的工具,它将基于OpenTelemetry的分布式追踪与故障注入相结合,适用于软件工程工作流中基于LLM的多智能体系统。该工具在工作流阶段、智能体步骤、智能体间通信、工具调用以及LLM调用中,对智能体执行进行与追踪对齐的遥测检测,并支持在选定交互点进行针对性故障注入。这使得能够以可复现的方式比较基线执行与故障执行,并通过对齐的追踪和运行工件检查效果。我们描述了该工具的动机、架构、实现、当前能力,以及在最小演示工作流和一个用于软件开发的真实基于LLM的多智能体系统上的初始验证情况。

英文摘要

Large Language Model-based multi-agent systems are increasingly explored for software engineering tasks, but they remain difficult to inspect, debug, and evaluate under controlled failures. We present llmmas-otel, a lightweight and framework-agnostic tool that combines OpenTelemetry-based distributed tracing with fault injection for LLM-based multi-agent systems in software engineering workflows. The tool instruments agent executions with trace-aligned telemetry across workflow phases, agent steps, inter-agent communication, tool calls, and LLM invocations, and supports targeted fault injection at selected interaction points. This makes it possible to compare baseline and faulty executions in a reproducible way and inspect the effects through aligned traces and run artifacts. We describe the motivation, architecture, implementation, current capabilities, and initial validation of the tool on a minimal demo workflow and a real LLM-based multi-agent system for software development.

DOI:10.1109/ICST69053.2026.00037

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑