发表机构
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院大学; 中国科学院自动化研究所; 中国科学院数学与系统科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对大型语言模型的过程规则推理不足,构建RuleWorld基准,提出DynaRule框架,在10K规则下将平均QA准确率提19个点,Recall@1超85%,大幅优于基线。
AI 中文摘要
大型语言模型(LLMs)擅长文本理解与生成,但在大规模可靠理解和应用外部提供的过程规则方面仍存在不足。为评估该能力,我们引入RuleWorld——一个大规模基准,其将规则重构为全局可复用的抽象单元,而非特定实例的事实。在RuleWorld中,设置了单规则、并行多规则及多跳推理等多个场景以开展全面评估。我们进一步提出DynaRule,这是一个端到端框架,能将给定规则注入键值(KV)缓存,并将检索转化为内部可学习的逐步过程。具体而言,DynaRule采用带有特殊<search>标记的堆叠步骤级注意力训练,以在推理过程中实现动态规则重注意力与更新。通过该方式,模型可在每一步重注意力于最相关的规则,动态替换过时规则,以支持更稳定的多步推理。在RuleWorld上的实验表明,现有LLMs在大规模规则库下面临挑战,而DynaRule将平均问答准确率提升了多达19个百分点,在10000条规则下实现了超过85%的Recall@1,大幅优于强基线模型。我们在此提供代码与数据集:this https URL。
英文摘要
Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWorld, a large-scale benchmark that reformulates rules as globally reusable abstract units rather than instance-specific facts. In RuleWorld, several scenarios, including single-rule, parallel multi-rule, and multi-hop reasoning, are settled for comprehensive evaluation. We further propose DynaRule, an end-to-end framework that injects the given rules into the KV cache and turns retrieval into an internal, learnable, step-wise process. Specifically, DynaRule employs Stacked Step-Level Attention Training with a special <search> token to enable dynamic rule re-attention and updating during inference. In this way, the model can re-attend to the most relevant rules at each step, dynamically replacing outdated ones to support more stable multi-step reasoning. Experiments on RuleWorld show that existing LLMs face challenges under large rule pools, while DynaRule improves average QA accuracy by up to 19 points and achieves over 85% Recall@1 at 10K rules, outperforming strong baselines by large margins. We make our code and dataset available here: https://github.com/SharkSpicy-NLP/Beyond-Factual-Knowledge.
CommentsAccepted by EMNLP 2026 Findings