LADDER:图引导的扩散语言模型用于高效多跳推理
LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning
浏览论文内容
中文总结 AI 辅助
LADDER通过图引导的并行解码将扩散语言模型与图检索增强生成结合,利用事件驱动自时钟检索和动态图传播,在多跳问答中提升准确率并大幅降低延迟。
中文摘要 AI 辅助
图检索增强生成(GraphRAG)通过利用结构化实体拓扑,显著增强了大型语言模型在复杂推理任务上的能力。然而,现有框架严重依赖标准自回归语言模型,其固有的顺序生成特性严重阻碍了整体推理效率。受扩散语言模型(DLMs)的启发,扩散语言模型通过连续并行细化解码提供大规模并行性,我们旨在离散空间中加速GraphRAG。然而,这面临两个非平凡的挑战。首先,部分去噪的草稿高度动态且不确定,使得动态图接地变得非平凡。其次,原始去噪状态本质上嘈杂且不稳定,使得同步图检索和多跳聚合在计算上代价高昂。为此,我们提出LADDER,一种新颖的框架,通过图引导的并行解码将扩散语言建模与GraphRAG桥接起来。具体而言,(i)我们提出一种事件驱动的自时钟检索,受我们关键洞察的启发:88%的目标实体在部分去噪状态早期出现,最终提交平均提前5.7-9.6步。该机制仅在可图链接实体集合扩展时动态触发图检索,产生一种异步自时钟策略,绕过学习门控或启发式阈值。(ii)设计了一个不完整查询图传播模块,使用专门的图基础模型处理新出现的实体查询,持续聚合多跳证据以锐化并行预测并加速整体解码收敛。在三个具有挑战性的多跳问答基准上的大量实验表明,LADDER将平均精确匹配从39.6%提升至45.2%,同时实现4.1倍的延迟降低。
英文摘要
Graph Retrieval-Augmented Generation (GraphRAG) has remarkably enhanced large language models on complex reasoning by leveraging structured entity topologies. However, existing frameworks heavily rely on standard autoregressive language models where the nature of inherent sequential generation severely hinders overall inference efficiency. Inspired by Diffusion Language Models (DLMs) that offer massive parallelism via continuous refine-in-parallel decoding, we aim to accelerate GraphRAG in the discrete space. However, it remains non-trivial for two challenges. First, partially denoised drafts are highly dynamic and uncertain, making dynamic graph grounding non-trivial. Second, raw denoising states are inherently noisy and unstable, making synchronous graph retrieval and multi-hop aggregation computationally prohibitive. To this end, we present LADDER, a novel framework that bridges diffusion language modeling with GraphRAG through graph-guided parallel decoding. Specifically, (i) we propose an event-driven self-clocking retrieval, inspired by our key insight that 88% of target entities emerge early in the partially denoised state, leading final commitment by an average of 5.7-9.6 steps. This mechanism dynamically triggers graph retrieval only when the set of graph-linkable entities expands, yielding an asynchronous self-clocking policy that bypasses learned gates or heuristic thresholds. (ii) An incomplete-query graph propagation module is designed to process the newly emerging entity queries using a specialized graph foundation model, continuously aggregating multi-hop evidence to sharpen parallel predictions and accelerate overall decoding convergence. Extensive experiments on three challenging multi-hop QA benchmarks show that LADDER raises average exact match from 39.6% to 45.2% while achieving a 4.1x latency reduction.
发表机构
- Huazhong University of Science and Technology(华中科技大学)
- Monash University(莫纳什大学)
- Tencent Youtu Lab(腾讯优图实验室)
机构由 AI 辅助整理,请以论文原文为准。