发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出通过加权DAG聚合集成LLM推理结构的框架,在六个基准上优于多数投票基线,可提供可检查的共识推理图,还能分析不同推理视角。
AI 中文摘要
大语言模型(LLMs)通过思维链探索问题,但这种探索被掩埋在非结构化文本中。在高风险任务中,用户无法判断哪些步骤有充分支撑、哪些替代方案被认真考虑过,也无法得知最终结论与模型舍弃的结论相比如何。我们提出一种框架,该框架集成多个LLMs的推理结构而非仅答案,方法是对从推理链中提取的有向无环图(DAG)进行加权合并。我们根据每个步骤被多少条推理轨迹独立验证来赋予权重,以返回“共识推理”。在涵盖法定解释、研究生水平科学、叙事多跳推理和一阶逻辑的六个基准测试中,我们的集成方法优于匹配预算的多数投票基线,在MuSR-MM(叙事多跳推理)上的最大准确率提升为3.1%。在单一模型上,该框架在相同推理轨迹预算下达到或超过自一致性方法的性能,同时还能提供可检查的共识推理图。集成权重与LLM评判者对推理质量的排名的斯皮尔曼相关系数ρ为0.30至0.51,在六个数据集中的五个数据集上,共识子图在两两比较中被偏好的比例为54.4%至65.4%,且我们发现该框架还可用于分析某个问题的不同推理视角。
英文摘要
Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously considered, or how the final conclusion compares to those the model discarded. We propose a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains. We weight each step by how many traces independently attest to it, to return "Consensus Reasoning". Across six benchmarks spanning statutory interpretation, graduate-level science, narrative multi-hop reasoning, and first-order logic, our ensemble outperforms a matched-budget majority-vote baseline, with a maximum accuracy gain of 3.1% on MuSR-MM (narrative multi-hop reasoning). On a single model, the framework matches or exceeds self-consistency at the same trace budget while additionally exposing an inspectable consensus reasoning graph. Ensemble weights correlate with LLM-judge rankings of reasoning quality at Spearman $ρ= 0.30$-$0.51$, and consensus subgraphs are preferred over alternatives leading to the majority-vote answer in 54.4-65.4% of head-to-head comparisons across five of six datasets. We observe that our framework can also be used to analyze diverse reasoning perspectives for a problem.