发表机构
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院; 中国科学院自动化研究所; 中国科学院大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有基准测试无法全面评估大型语言模型以规则为中心的场景推理能力的问题,本文提出RuleWeaver基准测试框架,通过构建规则化问答实例并开展过程级评估,发现当前主流LLMs在该任务上表现不佳。
AI 中文摘要
大型语言模型(LLMs)正越来越多地应用于专业领域,在这些领域中,有效利用领域专业知识往往需要对具体场景中的复杂规则进行推理。然而,现有的基准测试仅能部分评估这一能力,因为它们要么聚焦于输出层面的指令约束,要么忽略了规则在场景推理中所扮演的不同角色。为解决这些差距,本文提出了RuleWeaver,一个用于评估以规则为中心的场景推理的基准测试构建框架。RuleWeaver从语料库衍生的IF-THEN元规则出发,逐步将其扩充为复杂规则,并将这些规则组合成以规则为中心的场景问答实例。除了最终答案的正确性,RuleWeaver还支持基于评分标准的答案质量、规则召回率和规则准确率的过程级评估。对11个代表性LLMs的实验表明,当前模型在以规则为中心的复杂场景推理方面仍存在困难,即便表现最佳的模型也仅达到评分标准最大分数的约50%。我们在此处提供代码和数据集:this https URL。
英文摘要
Large language models (LLMs) are increasingly applied to specialized domains, where effective use of domain expertise often requires reasoning over complex rules in concrete scenarios. However, existing benchmarks only partially evaluate this capability, as they either focus on output-level instruction constraints or overlook the distinct roles that rules play in scenario reasoning. To address these gaps, this paper introduces RuleWeaver, a benchmark construction framework for evaluating rule-centered scenario reasoning. RuleWeaver starts from corpus-derived IF-THEN Meta Rules, progressively augments them into complex rules, and composes these rules into rule-centered scenario QA instances. Beyond final-answer correctness, RuleWeaver further supports process-level evaluation through rubric-based answer quality, rule recall, and rule precision. Experiments on 11 representative LLMs show that current models still struggle with complex rule-centered scenario reasoning, with even the best-performing model achieving only around 50% of the maximum rubric score. We make our code and dataset available here: https://github.com/SharkSpicy-NLP/RuleWeaver.
CommentsAccepted by EMNLP 2026 Findings