这是决策关键证据吗?学习验证规则治理决策
Is This Evidence Decision-Critical? Learning to Verify Rule-Governed Decisions
浏览论文内容
中文总结 AI 辅助
提出InterPact框架,通过反事实干预学习验证规则治理决策中的证据关键性,在单案例验证任务中达到68.28%准确率,优于六个基线。
中文摘要 AI 辅助
基于规则的推理,如资格检查和合同审查,要求语言模型根据个别条件评估证据,并在明确规则下组合其判断。证据评估中的错误可能使决策保持不变,但误解或忽略决策关键证据则可能逆转决策。识别此类证据使更强大的模型能够专注于检查相应的条件判断,支持准确和安全的决策。识别证据的关键性需要理解证据如何影响条件判断以及该判断如何影响决策。为实现这一目标,我们提出了一种基于干预的影响学习框架(InterPact),该框架能够在规则治理决策中对证据关键性进行反事实验证。具体而言,其证据干预构造器通过使用冻结语言模型编辑案件事实,同时保持规则和非目标条件不变,为传播验证器生成训练对。人工审核的标签记录由此产生的条件和决策变化,而完整的从状态到决策的映射则监督超出观察编辑的后果。在训练期间,验证器通过固定组合操作,将基于证据的条件概率加权到学习的条件决策预测中,将决策变化监督传播到基础模型。在推理时,训练好的基础模型直接根据原始案件和目标证据判断关键性,无需人工或更强模型的监督。在适应规则治理决策案例的单案例证据关键性验证中,InterPact达到了68.28%的准确率,优于所有六个基线。这些结果支持学习到的决策敏感性作为优先进行证据检查的基础。
英文摘要
Rule-based reasoning, as in eligibility checks and contract reviews, requires language models to assess evidence against individual conditions and combine their judgments under explicit rules. Errors in evidence assessment can leave a decision unchanged, but misinterpreting or overlooking decision-critical evidence can reverse it. Identifying such evidence allows more capable models to focus on checking the corresponding condition judgments, supporting accurate and safe decisions. Recognizing the evidence's criticality requires understanding how evidence affects a condition judgment and how that judgment affects the decision. To achieve the goal, we propose a INTERvention-based imPACT learning framework (InterPact), which enables counterfactual verification of evidence criticality in rule-governed decisions. Specifically, its evidence intervention constructor generates training pairs for a propagation verifier by editing case facts with a frozen language model while holding rules and non-target conditions fixed. Human-reviewed labels record the resulting condition and decision changes, while complete state-to-decision mappings supervise consequences beyond the observed edit. During training, the verifier weights learned conditional decision predictions by evidence-based condition probabilities through a fixed composition operation, propagating decision-change supervision into the base model. At inference, the trained base model directly judges criticality from the original case and target evidence, without human or stronger-model supervision. On single-case evidence criticality verification over adapted rule-governed decision cases, InterPact achieves 68.28% accuracy, outperforming all six baselines. These results support learned decision sensitivity as a basis for prioritizing evidence checks.
发表机构
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。