AI 中文总结
研究发现大语言模型推理能力蒸馏中,答案条件思维链会降低训练数据质量,导致可验证推理准确性下降,损失随难度增加,机制是从答案反向合理化,危害是数据属性,建议盲目生成答案。
AI 中文摘要
大语言模型推理能力蒸馏的标准方法是从模型中采样思维链,保留得出正确最终答案的那些,并对幸存者进行微调。当采样失败时,常见修复方法是向生成器展示正确答案并要求其写出达成该答案的思维链。我们表明第二步会以正确性过滤无法捕捉的方式降低训练数据质量。通过控制实验发现,在答案条件下生成思维链会大幅降低可验证推理准确性,损失随难度增加,在最难竞赛问题上高达约27分。机制在思维链中清晰可见,即从展示的答案反向合理化而非推导答案,早期的最终答案陈述是可衡量症状。危害是数据的属性而非生成器的,在任何微调前从无标签生成中就能看出,可在四个家族的八个思维模型中排序惩罚并跨教师家族转移。提示消融将其定位到朝着答案合理化的指令而非答案的单纯可见性。实际建议是盲目生成答案,因为没有正确性过滤能看到数据中的这种损害。
英文摘要
A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fails, a common fix shows the generator the gold answer and asks it to write a chain that reaches that answer. We show that this second step degrades the training data in a way that correctness filtering cannot catch. We run a controlled experiment that fixes the generator, the problem set, and the correctness filter, and varies only whether the chain is generated under answer-conditioning, the gold answer shown with a request to reach it. Training a strong instruction-tuned reasoning model on its own answer-conditioned chains sharply lowers its verifiable-reasoning accuracy. The loss grows with difficulty, reaching as much as about 27 points on the hardest competition problems. The mechanism is legible in the chains themselves, which rationalize backward from the shown answer instead of deriving it, with the early final-answer statement as the measurable symptom. The harm is a property of the data rather than the generator, read off unlabeled generations before any fine-tuning, ordering the penalty across eight thinking models from four families, and transferring across teacher families. A prompt ablation localizes it to the rationalize-toward instruction rather than the answer's bare visibility. The practical takeaway is to generate answer-blind, because no correctness filter can see this damage in the data.
Comments13 pages, 4 figures, 14 tables. Code: https://github.com/js-lee-AI/answer-leakage