arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非对称动态路由:平衡超图RAG中的推理深度与计算效率

Asymmetric Dynamic Routing: Balancing Reasoning Depth and Computational Efficiency in Hypergraph RAG

Qi Sun, Yijia Zhang, Caibo Li, Qiang Li, Yu Guo

arXiv 2609.29282首次发表:更新:

发表机构

School of Software Engineering, Xi'an Jiaotong University; EHV Power Transmission Company of China Southern Power Grid Co., Ltd.(西安交通大学软件学院; 中国南方电网有限责任公司超高压输电公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有RAG静态遍历策略在简单查询上冗余、复杂查询上不足的问题,提出非对称动态路由(ADR),通过轻量级分类器在三种拓扑遍历算子间动态调度,在保持推理性能的同时降低令牌消耗48.7%和延迟45.3%。

AI 中文摘要

尽管基于图和基于超图的检索增强生成(RAG)显著缓解了大语言模型(LLM)中的幻觉问题,但现有的基于结构的RAG系统通常采用静态遍历策略,而不考虑查询的复杂性。我们将这种“静态检索谬误”识别为简单查询计算冗余和复杂推理任务认知上下文缺口的主要来源。为了平衡推理质量与推理效率,我们提出了非对称动态路由(ADR),一种在分层知识图谱上运行的意图条件检索框架。ADR采用轻量级结构化分类器,在三种非对称拓扑遍历算子之间动态调度查询:局部事实锚定、自底向上邻接扩散和自顶向下洞察接地,这些算子共同实现跨分层知识层的双向信息流。跨五个领域特定语料库的大量实证评估表明,ADR在保持强大推理性能的同时,将提示词令牌消耗减少高达48.7%,端到端查询延迟减少45.3%,为查询自适应的超图RAG提供了有利的质量-效率权衡。

英文摘要

While graph-based and hypergraph-based Retrieval-Augmented Generation (RAG) significantly mitigate hallucinations in Large Language Models (LLMs), existing structure-based RAG systems typically adopt static traversal strategies regardless of the query complexity. We identify this ``static retrieval fallacy'' as a primary source of computational redundancy for simple queries and cognitive context gaps for complex reasoning tasks. To balance reasoning quality and inference efficiency, we propose Asymmetric Dynamic Routing (ADR), an intent-conditioned retrieval framework operating over hierarchical knowledge graphs. ADR employs a lightweight structured classifier to dynamically dispatch queries among three asymmetric topological traversal operators: localized fact anchoring, bottom-up adjacency diffusion, and top-down insight grounding, which collectively enable bidirectional information flow across hierarchical knowledge layers. Extensive empirical evaluations across five domain-specific corpora demonstrate that ADR maintains strong reasoning performance while reducing prompt token consumption by up to 48.7\% and end-to-end query latency by 45.3\%, yielding a favorable quality--efficiency trade-off for query-adaptive Hypergraph RAG.

Comments5 pages, 1 figures. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑