arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiagEvo:基于分层错误记忆的诊断引导自进化

Learning What to Practice: Diagnosis-Guided Self-Evolution for Language Models

Xincheng Wei, Yifan Ding, Fucheng Xiong, Yoshua Li, Dongsheng Ma, Rongxiang Weng, Xunliang Cai, Wenjian Ding, Yao Zhang

arXiv 2609.00768首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; Meituan; Peking University; Juntendo University(香港中文大学(深圳); 美团; 北京大学; 顺天堂大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DiagEvo利用求解器自身失败历史构建分层错误记忆,结合双置信度过滤实现无外部资源的自进化,在多基准测试中显著提升了Qwen系列等求解器的推理准确率。

AI 中文摘要

自博弈是语言模型自进化的有效范式,但缺乏引导时,求解器性能会在多轮迭代中趋于平稳或下降。无引导方法通过难度、可学习性或多样性等信号引导问题生成,这些信号能保持问题的挑战性和多样性,但未明确后续轮次应针对哪些未解决的推理弱点。有引导方法从外部任务资源(包括人类示例、文档语料库或指定难度目标)获取方向,因此依赖自博弈循环之外提供的任务信息。本文表明,所需方向可从求解器自身的失败历史中推导。我们提出DiagEvo,其诊断器从该历史中提取反复出现的错误原因,并将其存储在分层错误原因记忆中。该记忆将相关原因归为技能节点,并根据目标问题上的自一致性将每个节点标记为“活跃”或“已掌握”。挑战者利用这些状态和重复计数,在原因针对性生成与自由探索之间取得平衡。双置信度过滤仅在最常见的求解器答案具有明确投票优势时保留中等难度问题。DiagEvo从自博弈过程中产生的信息中推导其课程,无需外部任务资源。使用默认4B规模的诊断器时,DiagEvo在Qwen3-4B、Qwen3-8B和OctoThinker-8B这三个求解器的全部9个基准测试中,平均准确率均优于所有基线。在Qwen3-8B上,其在5个数学推理基准测试中的平均准确率达到72.3%,比R-Zero高出4.5个百分点;在全部9个基准测试中的平均准确率为57.4%,比DARC高出1.1个百分点。 ablation实验表明,分层错误原因记忆和双置信度过滤均对这些提升有贡献。

英文摘要

Self-play supports the self-evolution of language models, but solver performance can plateau or decline across rounds without guidance. Existing unguided methods typically use difficulty, learnability, or diversity signals to keep questions challenging and varied, without identifying which unresolved reasoning weaknesses to target. Existing guided methods rely on external task resources such as human examples, document corpora, or specified difficulty targets. We introduce DiagEvo, which guides question generation using the solver's failure history from self-play, without external task resources. Its diagnostician extracts recurring error causes and stores them in an error-cause memory. The memory groups related causes under skill nodes and tracks each as Active or Mastered according to self-consistency on targeted questions. The challenger uses these states and recurrence counts to balance cause-targeted generation with free exploration. Double-confidence filtering retains intermediate-difficulty questions only when the most common solver answer has a clear vote lead. With the default 4B diagnostician, DiagEvo outperforms all baselines in mean accuracy across nine benchmarks for each solver: Qwen3-4B, Qwen3-8B, and OctoThinker-8B. On Qwen3-8B, DiagEvo reaches 72.3% mean accuracy across five mathematical reasoning benchmarks, 4.5 percentage points above R-Zero. Its overall mean accuracy across nine benchmarks is 57.4%, 3.5 percentage points above SPICE. Ablations show that mixed generation, memory-state updates with cross-state stitching, and double-confidence filtering contribute to these gains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑