Loom:通过嵌入空间重加权将诊断线索编织为自由文本共识
Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
浏览论文内容
中文总结 AI 辅助
本文提出用于根因分析的Loom框架,通过嵌入空间重加权聚合模块化启发式的开放式假设,在OpenRCA基准上实现了准确率与效率的帕累托优化,为工业场景提供可靠共识方案。
中文摘要 AI 辅助
在工业界实际部署自然语言处理(NLP)系统时,将嘈杂、冲突的文本假设聚合为可靠共识是一项基本挑战。整体式大语言模型(LLM)智能体为根因分析(RCA)等任务提供了无限表现力,但存在上下文限制、幻觉累积和推理延迟过高的问题;传统弱监督虽具备统计严谨性,但数学上仅适用于离散类别。本文提出Loom,一款用于实际根因分析(RCA)的生成式共识框架,它桥接了上述两种范式。Loom将模块化启发式方法生成的开放式假设(动态填充特定事件实体、时间和指标的诊断模板)投影到连续嵌入空间中,通过迭代基于质心的重加权算法解决冲突信号,所得共识权重用于单次轻量LLM合成步骤。在OpenRCA基准上评估显示,Loom处于准确率-效率帕累托前沿:它在Bank和Market-2数据集上与最先进的自主智能体表现相当,在Market-1和Telecom数据集上稍逊,且在所有四个数据集上每事件仅需1次LLM调用(速度提升约26倍;若使用80亿参数合成器则提升约33倍)。本文还讨论了部署经验,强调了智能体深度与推理延迟之间的权衡、冗余检测的负面结果,以及确定性共识如何提升领域专家(SME)的信任度。
英文摘要
Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge when deploying NLP systems in real-world industrial settings. While monolithic Large Language Model (LLM) agents offer unbounded expressivity for tasks like Root Cause Analysis (RCA), they suffer from context limits, compounding hallucinations, and prohibitive inference latency. Traditional weak supervision offers statistical rigor but is mathematically restricted to discrete classes. We present Loom, a generative consensus framework deployed for real-world RCA that bridges these paradigms. Loom aggregates open-form hypotheses emitted by modular heuristics (diagnostic templates dynamically populated with episode-specific entities, times, and metrics) by projecting them into a continuous embedding space, and resolves conflicting signals with an iterative centroid-based reweighting algorithm. The resulting consensus weights ground a single lightweight LLM synthesis step. Evaluated on the OpenRCA benchmark, Loom occupies the accuracy--efficiency Pareto frontier: it matches a state-of-the-art autonomous agent on Bank and Market-2 and trails on Market-1 and Telecom, while using a single LLM call per incident on all four datasets ($\sim$26$\times$ faster; $\sim$33$\times$ with an 8B-parameter synthesizer). We discuss our deployment experience, highlighting lessons learned regarding the trade-offs between agentic depth and inference latency, negative results in redundancy detection, and how deterministic consensus fosters trust among Subject Matter Experts~(SMEs).
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。