arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03674cs.LG

DiagLoop:面向诊断大语言模型的带阶段局部强化的反事实数据飞轮

DiagLoop: A Counterfactual Data Flywheel with Stage-Localized Reinforcement for Diagnostic LLMs

Jian Zhang, Bingyi Wang, Yizhi Liu

AI总结:

本文提出DiagLoop反事实数据飞轮,利用合成场景训练8B诊断大语言模型,在工业系统与疾病类别上的严格路径正确性及对照性能均优于传统基线与专有参考模型。

AI中文摘要:

因果诊断模型必须解释结论如何从证据中推导得出,因为诊断结果指导修复与治疗。然而,严重病例稀缺、记录极少包含推理路径、数据难以跨配置迁移,这些问题为本地部署带来了挑战。本文提出DiagLoop,一种反事实数据飞轮,它将针对每个机制家族一次性编写的编码物理关系或临床指南,转化为超越已记录病例的训练监督信号。仅用于训练的教师模型通过改变原因、上下文和观测值来生成反事实世界,而独立的混合检查器仅接纳有效世界。学生模型通过症状抽象、因果链构建和根本原因归因进行推理。阶段特定标准可识别学生模型的最早失效点:对于非终端失效,有界修复会探测下游能力,生成的弱点轮廓会指导后续数据生成;阶段局部强化学习仅更新模型生成的延续内容,重放与保存机制可减少遗忘。相同标准通过与提议者分离的检查,用于接纳、归因、奖励与再生流程。仅使用合成场景、无需病例级专家推理注释的情况下,所得到的8B模型在严格路径正确性上优于最强传统基线:在8个工业系统中提升11.6个百分点,在10个疾病类别中提升5.5个百分点;在与混乱路由对照的比较中,分别提升3.9和2.3个百分点;该模型在两个领域均超过评估的专有参考模型,即便这些参考模型仅获得少样本示例或上下文规格信息。

英文摘要:

Causal diagnostic models must explain how conclusions follow from evidence because diagnoses guide repairs and treatments. Yet serious cases are scarce, records rarely contain reasoning paths, and data transfer poorly across configurations, complicating local deployment. We present DiagLoop, a counterfactual data flywheel that converts codified physical relations or clinical guidelines, authored once per mechanism family, into training supervision beyond recorded cases. A training-only teacher proposes counterfactual worlds by varying causes, contexts, and observations, while an independent hybrid checker admits only valid worlds. The student reasons through symptom abstraction, causal-chain construction, and root-cause attribution. Stage-specific criteria identify its earliest failure. For nonterminal failures, a bounded repair probes downstream competence, and the resulting weakness profile guides subsequent data generation. Stage-localized reinforcement learning updates only the model-generated continuation, while replay and preservation reduce forgetting. The same criteria govern admission, attribution, reward, and regeneration through checks separate from the proposer. Using only synthesized scenarios and no case-level expert reasoning annotations, the resulting 8B model improves strict path correctness over the strongest conventional baseline. Gains are 11.6 points across eight industrial systems and 5.5 points across ten disease categories. Gains over a deranged-routing control are 3.9 and 2.3 points, respectively. The model also exceeds the evaluated proprietary references in both domains, even when they receive few-shot examples or the specification in context.

补充信息

↑