用于证据推理的反事实自进化智能体
Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning
浏览论文内容
中文总结 AI 辅助
提出反事实自进化框架,通过生成反事实上下文为冻结求解者提供证据,以提升证据推理能力,在临床、事实核查和商业推理任务上超越前沿模型。
中文摘要 AI 辅助
自我对弈的提出者-求解者方法通过生成任务并从已验证的解决方案中学习来提升推理能力。然而,对于证据可识别任务(其中特定案例的证据和领域知识决定了一个可核查的答案),自我对弈需要生成其答案可以被独立验证的合理案例。我们引入了反事实自进化,它生成反事实上下文以重新考虑原始案例。一个可训练的提出者构建有针对性的证据编辑,并通过因果解释描述潜在的结果变化。我们手工制作了一个专家验证的反事实指令微调数据集,以教会提出者在广泛的行动-结果场景中生成高质量的反事实。每个反事实指令微调示例都指定了定义类别内的编辑,并解释了其对决策的假设性因果效应,教会提出者系统地推理什么发生了变化以及为什么。我们在这些示例上对提出者进行指令微调,然后制定一个整合求解者和验证者反馈的微调奖励。在多样化的反事实场景中,该奖励偏向于高质量的反事实和合理的修订,同时惩罚推翻正确决策的更改。反事实上下文旨在纠正错误并加强对正确决策的信心。被接受的反事实累积在记忆中,为冻结的求解者提供上下文证据;求解者通过演化的上下文而非权重更新来适应。我们将该框架应用于临床推理、事实核查和商业推理。我们的评估跟踪了随着反事实记忆增长,在连续轮次中的性能,包括向更难案例的迁移。我们的方法在多种前沿模型上取得了优越的结果。
英文摘要
Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action--outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.