SKIP:一个自知识引导的逐步偏好学习框架,用于简洁推理
SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning
- Beijing University of Posts and Telecommunications(北京邮电大学)
- QuanCheng Laboratory(泉城实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SKIP框架通过自知识引导的逐步偏好学习,结合轻量级微调和DPO,在压缩推理长度的同时保持准确性,并展现出良好的泛化能力。
AI中文摘要:
虽然链式思维(CoT)推理已被证明是有效的,但它常常导致过度思考,从而在大语言模型(LLMs)中引发计算开销、推理延迟,甚至性能下降。现有的简洁推理框架在压缩输出长度的同时显著牺牲了准确性。在本文中,我们提出了SKIP,一个自知识引导的逐步偏好学习框架。首先通过轻量级微调来调整模型的输出风格,SKIP引入了一种精心设计的知识探测机制,引导模型在每个推理步骤输出答案。基于中间步骤的正确性,我们构建偏好数据,利用DPO引导模型走向更高效且正确的推理。实验结果表明,我们的方法在微调后有效提升了推理压缩,同时缓解了性能下降。此外,SKIP在分布外数据集上展现出强大的泛化能力。我们进一步对我们的框架的组件参数进行了消融研究。
英文摘要:
While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference latency, and even degraded performance in large language models (LLMs). Existing concise reasoning frameworks significantly compromise accuracy while compressing the length of output. In this paper, we propose SKIP, a self-knowledge-guided step-wise preference learning framework. Starting with lightweight fine-tuning to adjust the model's output style, SKIP introduces a carefully designed knowledge probing mechanism to guide model to output an answer at each reasoning step. Based on the correctness of intermediate steps, we construct preference data that guide the model toward more efficient and correct reasoning by leveraging DPO. Experimental results demonstrate that our method effectively improves reasoning compression while mitigating performance degradation after fine-tuning. Besides, SKIP shows strong generalization ability on out-of-distribution datasets. We further conducted ablation studies on the component parameters of our framework.