Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
通过小规模偏好优化修剪大推理模型的长推理链
机构 * State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) ; University of Science and Technology of China(中国科学技术大学) ; Institute of Artificial Intelligence(人工智能研究院) ; Hefei Comprehensive National Science Center(合肥综合性国家科学中心) ; Central China Normal University(中部师范大学) ; Meituan(美团) ; NeoShell
专题命中 推理与问题求解 :preference optimization(title,abstract);分类 cs.AI
AI总结 本文提出LCPO方法,通过统一的Bradley-Terry损失框架,有效减少大推理模型的输出长度,同时保持推理性能,实验显示在多个基准上平均输出长度减少超过50%。
Comments 26 pages, 8 figures, ICLR'26