TriCalRAG:一种用于AIOps中基于本地部署LLM的根因分析的三策略检索增强基准
TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps
浏览论文内容
中文总结 AI 辅助
TriCalRAG基准评估本地部署的开放权重LLM在AIOps根因分析中的表现,证明RAG策略优于零样本提示,提高F1分数并稳定模型校准,同时提供吞吐量和量化优化分析。
中文摘要 AI 辅助
云托管的大型语言模型(LLM)越来越多地用于AIOps流水线中的根因分析(RCA),但它们引入了数据隐私风险、网络延迟和每次查询成本,这些成本随生产日志量的增加而难以扩展。我们提出了TriCalRAG,一个基准测试,用于评估通过vLLM在单个高内存工作站GPU(NVIDIA RTX PRO 6000,96GB)上本地服务的开放权重LLM,并与基于LSTM的经典日志异常检测器(DeepLog)进行比较,使用四个真实、公开可用的日志数据集(BGL、HDFS、Thunderbird、OpenStack)。我们在三种提示策略下评估了两个开放权重模型(Qwen2.5-14B、Mistral-Small):零样本、少样本和基于标记事件历史的检索增强生成(RAG),报告了准确率、精确率/召回率和F1分数,以及基于3个随机种子的bootstrap 95%置信区间,同时报告了吞吐量和VRAM占用。我们的结果表明,RAG不仅将平均F1比零样本提示提高了0.10-0.27,更重要的是,显著稳定了模型校准:零样本提示使两个模型趋向于接近退化的行为(在某些数据集上预测“异常”高达100%的事件),而RAG在大多数配置中将预测阳性率保持在接近真实类别平衡的水平。Mistral-Small的宏平均F1高于Qwen2.5-14B(0.644对比0.560),但在更多配置中出现校准失败(12个中的7个对比5个),同时吞吐量约为后者的一半——这表明更好的模型选择取决于部署是优先考虑峰值准确率还是跨提示条件的可预测行为。消融实验显示,批处理在单卡上将吞吐量扩展了41倍,4位量化将延迟降低了20%,且没有可测量的准确率损失。我们发布了我们的基准测试工具、数据集划分和评估代码,以支持可复现的本地AIOps研究。
英文摘要
Operational logs create a need for private, resource-efficient incident analysis, but aggregate detection scores can conceal severe prediction bias. We present TriCalRAG, a reproducible benchmark for log-anomaly detection with generated root-cause and remediation outputs across BGL, HDFS, Thunderbird, and OpenStack. The primary evaluation compares Qwen2.5-14B and Mistral-Small-22B, served through vLLM on one NVIDIA RTX PRO 6000 GPU (96 GB), under zero-shot, few-shot, and retrieval-augmented generation (RAG) prompting across three data-sampling seeds. We report F1, bootstrap confidence intervals, predicted-positive rates, throughput, and memory use, with DeepLog as a held-out classical baseline. Mistral-Small attains a higher macro-averaged F1 than Qwen2.5-14B (0.644 versus 0.560), whereas Qwen provides approximately twice the throughput. A separate log-probability evaluation compares raw decisions with Contextual Calibration (CC) and Batch Calibration (BC): neither correction consistently improves prediction-balance diagnostics across prompting strategies. Supplementary single-run comparisons extend evaluation to 4-bit Llama-3.1-70B via local Ollama and Claude Haiku 4.5 via Anthropic's cloud API; RAG improves F1 on all four datasets for both models. Claude's reported aggregate F1 increases from 0.566 to 0.695, while the estimated API cost rises from \$0.94 to \$2.15 per 1,000 incidents. Local deployment ablations show approximately 41-fold throughput scaling with batching and 20% lower latency with 4-bit quantization on the tested workload. These findings support retrieval as useful context for anomaly decisions, while the supplementary protocols, unvalidated explanation quality, and prediction-balance diagnostics limit broader claims about RCA accuracy and probabilistic calibration.
发表机构
- Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校)
机构由 AI 辅助整理,请以论文原文为准。