arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21911cs.SEcs.DL

利用已解决的事件历史进行大语言模型辅助的软件错误诊断

Leveraging Resolved Incident History for LLM-Assisted Software Bug Diagnosis

Boyuan Guan, Hailu Xu, Jamie Rogers

首次发表
浏览论文内容

中文总结 AI 辅助

研究利用已解决事件历史进行软件错误诊断,提出OM-RAG方法,将问题索引为三元组并单跳嵌入检索先例,为大语言模型管理员提供支持,实验表明该方法在诊断准确率和修复正确性上优于其他方法。

中文摘要 AI 辅助

有效的软件错误诊断需要正确的知识来源(操作失败历史,而非仅系统文档)和正确的检索结构(结构化记录,而非非结构化块)。当前检索增强生成(RAG)方法在一个或两个维度上存在不足。我们提出操作记忆RAG(OM-RAG),将已解决问题索引为结构化症状-根本原因-解决方案三元组,并通过单跳嵌入检索最相似的历史先例。OM-RAG为一个专用大语言模型管理员提供支持,该管理员已运行生产Dataverse实例超过六个月,利用系统自身已解决的失败历史实现监控-分析-计划-执行-知识(MAPE-K)反馈循环的知识组件。在一个包含四种配置的消融实验中,基于大语言模型的判断涵盖所有1172个问题,人工验证了一个123个问题的子集。OM-RAG实现了0.931的诊断准确率和0.809的修复正确性,优于基于块的检索(提高186%)和基于文档的概念图(提高77%)。无模型检索指标(平均相似度0.880)独立证实了排名。

英文摘要

Effective software bug diagnosis requires two ingredients: the right knowledge source (operational failure history, not just system documentation) and the right retrieval structure (structured records, not unstructured chunks). Current retrieval-augmented generation (RAG) approaches fall short on one or both dimensions. We propose Operational Memory RAG (OM-RAG), which indexes resolved issues as structured symptom-root cause-resolution triples and retrieves the most similar historical precedent via single-hop embedding. OM-RAG powers a purpose-built large language model (LLM) administrator that has operated a production Dataverse instance for over six months, operationalizing the Knowledge component of the Monitor-Analyze-Plan-Execute-Knowledge (MAPE-K) feedback loop with the system's own resolved failure history. In a controlled four-configuration ablation, LLM-based judging covers all 1,172 issues and manual verification a 123-issue subset. OM-RAG achieves Diagnosis Accuracy of 0.931 and Fix Correctness of 0.809, outperforming chunk-based retrieval (+186%) and a documentation-anchored concept graph (+77%). Model-free retrieval metrics (mean similarity 0.880) independently corroborate the ranking.

补充信息

↑