直到证据确凿:教LLM调查员何时结案
Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case
浏览论文内容
中文总结 AI 辅助
本研究针对LLM调查员的结案决策,提出三项测试并构建Nautil数据集,通过微调与强化学习使9B模型结案更依赖证据,高估率大幅下降,准确率显著提升。
中文摘要 AI 辅助
事故、缺陷和停机调查以普通问答从未面临的决策告终:即目前收集的证据是否足以结案。我们针对LLM调查员研究这一决策,他们从案件档案中请求证据,修订其假设,并要么基于所读内容以结论结案,要么保持案件未结并指出缺失之处。这一判断并非随能力而来:未经训练的9B模型在97%的回答中高估其证据,而能识别正确原因(在84%的案件中)的前沿模型仍有91%的高估,并在41个官方结论为“原因未定”的案件中关闭了17个。衡量这一判断也非易事:案件来源在很大程度上预测其标签,仅读取来源的规则在我们的测试案例上达到83.0的平衡准确率。因此,我们通过三项测试评估结案:结案准确率(对照此规则及每个来源内部报告)、证据依赖性(移除结论依据并检查模型是否停止结案)、以及结论与缺口质量(对模型所断言和所指出缺失之处的评判性检查清单)。我们构建了Nautil,包含来自航空、铁路、海事、化学品安全和车辆缺陷报告以及生产服务器事件的731个经审计案件,附有教师轨迹、分布外测试集和反事实证据版本。在这些轨迹上微调9B模型使其结案遵循证据:移除依据使其结案率相对于匹配对照组降低26个百分点,高估从97%降至35%,正确且未高估的结论从3%升至43%。仅奖励结案决策的强化学习随后将平衡准确率从69.2提升至83.3,与教师持平,来源内准确率从60.4提升至74.1,但证据依赖性有所牺牲。
英文摘要
Accident, defect and outage investigations end with a decision that ordinary question answering never faces: whether the evidence gathered so far is enough to close the case. We study this decision for LLM investigators, which request evidence from a case file, revise their hypotheses, and either close the case with a conclusion grounded in what they read or leave it open and name what is missing. This judgment does not come with capability: an untrained 9B model overstates its evidence in 97% of its answers, and a frontier model that identifies the right cause in 84% of cases still overstates in 91% and closes 17 of the 41 cases whose official finding is "cause undetermined". Measuring it is also non-trivial: the source of a case largely predicts its label, and a rule that reads only the source reaches 83.0 balanced accuracy on our test cases. We therefore evaluate closure with three tests: closure accuracy, reported against this rule and within each source; evidence dependence, which removes the grounds of a conclusion and checks whether the model stops closing; and conclusion and gap quality, a judged checklist of what the model asserts and what it says is missing. We build Nautil, 731 audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents, with teacher trajectories, an out-of-distribution test set and counterfactual evidence versions. Fine-tuning a 9B model on these trajectories makes its closures follow the evidence: removing the grounds lowers its closure rate by 26 points relative to a matched control, overstatement falls from 97% to 35%, and correct, non-overstated conclusions rise from 3% to 43%. Reinforcement learning that rewards only the closure decision then raises balanced accuracy from 69.2 to 83.3, on par with the teacher, and within-source accuracy from 60.4 to 74.1, at some cost in evidence dependence.
发表机构
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。