arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09587cs.LGcs.AIcs.CL

协作推理蒸馏:基于交叉反馈与连贯性策展

Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation

  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Taehoon Kim, Seunggeun Cho, Dongsu Han

AI总结:

提出协作推理蒸馏框架,通过交叉反馈、逐步质量评估和连贯步骤拼接,以极小数据训练出超越基线的紧凑推理模型CRD-4B。

AI中文摘要:

推理能力对于推进大型语言模型至关重要,然而当前的方法要么需要巨大的计算预算,要么难以有效地将推理能力蒸馏到较小的模型中。标准的蒸馏方法依赖于基于结果的奖励,无法区分合理的推理与幸运的猜测。我们提出了协作推理蒸馏(CRD),一个通过三项创新来增强紧凑模型推理能力的框架:(1)交互式交叉反馈,教师迭代地批评彼此的推理;(2)细粒度的逐步质量评估,捕获独立于最终答案的逻辑有效性;(3)连贯性感知的步骤拼接,综合互补的优势。学生模型通过带有预算约束的推理质量优化(RQO)进行训练。我们的模型CRD-4B在MATH-500上达到97.3%的准确率,在AIME'25上达到70.3%,超越了基线,同时仅使用5万个训练样本,比同类模型的数据集小至多12倍。

英文摘要:

Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and lucky guesses. We propose Collaborative Reasoning Distillation (CRD), a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other's reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths. Students are trained via Reasoning Quality Optimization (RQO) with budget constraints. Our model, CRD-4B, achieves 97.3% on MATH-500 and 70.3% on AIME'25, surpassing baselines while using only 50K training examples, up to 12 times smaller than the datasets of comparable models.

补充信息

↑