AI 中文总结
CausalOPD 是一种课程在线过程提炼框架,通过识别首错误步并结合短视界强化学习,提升了学生模型的因果链推理能力,在三个领域中表现优于序列级提炼及专有参考模型。
AI 中文摘要
临床诊断、法律判决、工业故障诊断等诸多关键推理任务依赖于依赖步骤的因果链,其中早期错误会传播,而正确结论可能掩盖无效推理。尽管大语言模型在这类任务上表现良好,但隐私、延迟和可控性方面的需求促使将其提炼为可本地部署的模型。标准轨迹模仿无法修正学生模型自身 rollout 分布上的过程错误。我们提出 CausalOPD,一种课程在线过程提炼框架。知识增强型教师首先提供基于特定领域因果规则、实体关系和结构约束的轨迹。学生随后生成策略内轨迹,教师识别首个错误步,即最早的可验证违反可用约束的转移。从已验证的前缀开始,短视界强化学习修复此局部失败。因果阶段课程按照传播顺序从证据级错误推进到机制级错误,再到结论级错误。在三个领域中,CausalOPD 相比序列级在线过程提炼将平均路径正确率提升了 23.4 个百分点,并将“标签正确但推理错误”的比率从 15.7% 降至 4.4%。特定领域的 8B 学生模型在所有领域的路径正确率上也超过了所有评估的专有参考模型。
英文摘要
Many critical reasoning tasks, including clinical diagnosis, legal judgment, and industrial fault diagnosis, require step-dependent causal chains in which early errors propagate and correct conclusions can mask invalid reasoning. Although large language models perform well on such tasks, privacy, latency, and controllability motivate distillation into locally deployable models. Standard trajectory imitation does not correct process errors on the student's own rollout distribution. We propose CausalOPD, a curriculum online process distillation framework. A knowledge-augmented teacher first provides trajectories grounded in domain-specific causal rules, entity relations, and structural constraints. The student then generates on-policy trajectories, and the teacher identifies the first wrong step, defined as the earliest transition that verifiably violates available constraints. Starting from the verified prefix, short-horizon reinforcement learning repairs this localized failure. A causal-stage curriculum advances from evidence-level to mechanism-level and conclusion-level errors, following their propagation order. Across three domains, CausalOPD improves average path correctness by 23.4 percentage points over sequence-level online process distillation and reduces the right-label-wrong-reasoning rate from 15.7% to 4.4%. The domain-specific 8B students also surpass both evaluated proprietary references in path correctness across all domains.
Comments9 pages, 2 figures