发表机构
Institute of Smart City Technologies, University of Macau; ByteDance(澳门大学智慧城市技术研究院; 字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对规则治理决策任务,提出RGDT-Bench基准,提供202.1K条件级监督,并训练带论证监督的奖励模型,将任务平均AUROC提升至69.24%,优于现有评估器。
AI 中文摘要
我们研究规则治理决策任务(RGDTs)中的推理,在此类任务中,模型将外部规则应用于案件事实并证明决策的合理性,这在政策、合同和合规场景中均有要求。除了标准数学和逻辑推理任务所强调的演绎能力外,RGDTs还要求解释规则及其适用性、根据证据评估条件、在规则和例外情况下综合判断,并提供可核查的论证理由。这些需求催生了一个同时评估决策及其陈述理由的基准。我们引入RGDT-Bench,在四个任务轨道和八种支持的任务-探针组合中提供202.1K个条件级监督槽位,这些组合在对支持信息的访问上有所不同。标签盲提取和确定性检查生成论证完整性的标签:源引用覆盖率和陈述决策理由的一致性。该基准将失败归因于四个处理层:规则使用、条件、证据和聚合,并检查最终结果。在可评估的正确响应中,六个评估的LLM和支持的任务-探针组合中,论证不完整性平均为40.2%。这种论证不完整性构成潜在的安全风险,且难以检测:十七个现有评估器中最好的仅达到57.69%(随机:50%)的任务平均接收者操作特征曲线下面积(AUROC)。为解决这一难题,我们训练了一个带有论证监督的简单奖励模型。它在正确答案中达到69.24%的任务平均AUROC,超过匹配的结果监督基线10.37个百分点(pp),超过最佳现有评估器11.55个百分点。除了完整性评估外,该模型在几乎所有响应选择比较中均优于结果监督基线,支持RGDT-Bench的论证监督用于RGDT推理。
英文摘要
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require interpreting rules and their applicability, assessing conditions from evidence, combining judgments under rules and exceptions, and providing checkable justifications. These demands motivate a benchmark assessing both decisions and their stated grounds. We introduce RGDT-Bench, providing 202.1K condition-level supervision slots across four task tracks and eight supported task-probe combinations that vary access to supporting information. Label-blind extraction and deterministic checks produce labels for warrant completeness: source-referenced coverage and consistency of stated decision grounds. The benchmark attributes failures to four process layers: rule use, condition, evidence, and aggregation, and checks the final outcome. Among evaluable correct responses, warrant incompleteness averages 40.2% across six evaluated LLMs and supported task-probe combinations. Such warrant incompleteness poses potential safety risks and remains difficult to detect: the best of seventeen existing evaluators reaches only 57.69% (random: 50%) task-averaged area under the receiver operating characteristic curve (AUROC). To address this difficulty, we train a simple reward model with warrant supervision. It achieves 69.24% task-averaged AUROC among correct answers, exceeding the matched outcome-supervised baseline by 10.37 pp (percentage points) and the best existing evaluator by 11.55 pp. Beyond completeness assessment, the model outperforms both outcome-supervised baselines across nearly all response-selection comparisons, supporting RGDT-Bench's warrant supervision for RGDT reasoning.
Comments33 pages, 12 figures, 20 tables