arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25583cs.CL

GRIP:用于高效推理的粒度奖励引导参数插值

GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning

Lam So, Canhui Wu, Han Lin

首次发表
浏览论文内容

中文总结 AI 辅助

针对推理模型与指令模型的准确性-效率不匹配问题,提出GRIP框架,通过优化同架构两模型模块的可学习插值比例,实现更优的准确性-效率权衡。

中文摘要 AI 辅助

面向推理的大型语言模型通常通过生成长思维链实现强大的问题解决性能,但该行为会大幅提升推理成本与延迟。相比之下,指令微调模型的回答更简洁,但往往缺乏相当的推理能力。这种准确性与效率的不匹配促使我们提出一种轻量方法,无需完整模型重训练即可结合两类模型的优势。本文提出GRIP(Granular Reward-guided Interpolation of Parameters,粒度奖励引导参数插值),这是一种用于高效推理的奖励引导参数插值框架。给定架构相同的推理模型与指令模型,GRIP为各个模块分配可学习的插值比例,仅优化这些比例,同时保持两个源模型冻结。插值比例通过奖励信号训练,该信号偏向既正确又简洁的回答。实验表明,GRIP比固定或基于搜索的合并基准实现了更优的准确性-效率权衡,还揭示了与高效推理相关的模块级融合模式。

英文摘要

Reasoning-oriented large language models often achieve strong problem-solving performance by generating long chains of thought, but this behavior substantially increases inference cost and latency. In contrast, instruction-tuned models tend to answer more concisely, yet often lack comparable reasoning ability. This accuracy-efficiency mismatch motivates a lightweight approach that combines the strengths of both models without full model retraining. In this paper, we propose GRIP (Granular Reward-guided Interpolation of Parameters), a reward-guided parameter interpolation framework for efficient reasoning. Given a reasoning model and an instruction model with identical architectures, GRIP assigns learnable interpolation ratios to individual modules and optimizes only these ratios while keeping both source models frozen. The interpolation ratios are trained with a reward signal that favors responses that are both correct and concise. Experiments show that GRIP achieves a better accuracy-efficiency trade-off than fixed or search-based merging baselines and further reveals module-wise fusion patterns associated with efficient reasoning.

发表机构

  • Peking University(北京大学)
  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑