arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26094cs.LG

CoEvo:单一模型的预言机接地自进化用于多步因果推理

CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning

Jian Zhang, Bingyi Wang, Yizhi Liu

首次发表
浏览论文内容

中文总结 AI 辅助

CoEvo利用预言机验证单步,通过提议者-求解者角色交替实现单一模型自进化,在多步因果推理基准上超越蒸馏基线,路径正确率达82.1%。

中文摘要 AI 辅助

多步因果推理需要链式推理,其中每一步都约束下一步。早期错误会无声传播,而通过有缺陷的逻辑得出的正确答案会逃避结果层面的检测。在专业领域,教师大语言模型在中间步骤上出错,安全约束限制了云端蒸馏,且变化的条件要求适应性,这使得自进化成为切实可行的途径。朴素的自进化可能会崩溃:仅基于结果的奖励让模型利用分布捷径,而薄弱的自我评估将虚假路径强化为稳定的失败模式。我们利用一个关键的不对称性:生成正确的链条很难,但验证单一步骤很容易。许多高风险领域允许存在确定性的、可查询的预言机,即物理模拟器或基于编码化约束的规则引擎。它无需教师级能力即可检查断言的步骤,并在其规则之外弃权(不执行);它能检查模型所断言的内容,但绝不替代模型。这促成了CoEvo,一个预言机接地的自进化框架,其中单一模型在提议者与求解者之间交替。作为求解者,模型生成竞争性链条;组内辩论暴露分歧步骤,作为能力边界的代理,预言机将其裁决为过程级监督。作为提议者,同一模型在预言机约束内构建逐渐更难的场景,将课程导向深层多跳链条。两种角色联合更新,因此训练压力与模型共同进化。在工业、临床和法律多步因果推理基准上,CoEvo使8B大语言模型维持自进化,在路径正确性上超越蒸馏基线和最强专有参考(82.1%对71.4%)。训练后的模型泛化到未见类别和系统,保持根因准确性。

英文摘要

Multi-step causal reasoning requires chaining inferences where each step constrains the next. An early error propagates silently, and a correct answer reached via flawed logic evades outcome-level detection. In specialized domains, teacher LLMs err on intermediate steps, safety constraints restrict cloud distillation, and shifting conditions demand adaptation, leaving self-evolution as the practical route. Naive self-evolution can collapse: outcome-only rewards let the model exploit distributional shortcuts, and weak self-evaluation reinforces spurious paths into stable failure patterns. We exploit a key asymmetry: generating a correct chain is hard, but verifying a single step is easy. Many high-stakes domains admit a deterministic, queryable oracle, a physics simulator or rule engine over codified constraints. It checks asserted steps without teacher-level ability and abstains beyond its rules; it can check what the model asserts, never replace it. This enables CoEvo, an oracle-grounded self-evolution framework where a single model alternates between Proposer and Solver. As Solver, the model generates competing chains; intra-group debate exposes disagreement steps, a proxy for the capability boundary, and the oracle adjudicates them into process-level supervision. As Proposer, the same model constructs progressively harder scenarios inside oracle constraints, steering the curriculum toward deep multi-hop chains. Both roles are updated jointly, so training pressure co-evolves with the model. On industrial, clinical, and legal multi-step causal reasoning benchmarks, CoEvo enables an 8B LLM to sustain self-evolution, surpassing distillation baselines and the strongest proprietary reference on path correctness (82.1% vs. 71.4%). The trained model generalizes to unseen categories and systems, preserving root-cause accuracy.

补充信息

↑