arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39346cs.CL

离线指导,在线推理:为小型语言模型重用LLM反馈

Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models

Bohan Zhang, Linan Yue, Weibo Gao, Pengyu Chen, Hong Guo, Yanqi Hao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出可重用潜在修正(RLC),在仅离线访问LLM且SLM参数固定的约束下,将LLM反馈转化为SLM隐藏空间中的持久经验,实现无需在线LLM调用的推理增强,实验验证其有效性。

中文摘要 AI 辅助

大型语言模型(LLMs)提供强大的推理能力,但通过商业API访问通常成本高昂,而小型语言模型(SLMs)更易于本地部署,但推理能力较弱。这种能力与部署之间的差距促使了LLM-SLM协作,旨在利用LLM的能力提高SLM推理,同时保留SLM的部署优势。现有方法主要遵循两种范式。知识蒸馏使用LLM生成的答案和推理轨迹离线训练SLMs,但需要参数更新和额外训练。或者,在线协作将困难问题路由到LLM,或在SLM遇到困难时利用LLM生成的指导和修正。尽管有效,在线协作需要重复访问LLM。此外,为特定问题生成的指导在推理后被丢弃,无法惠及涉及类似推理状态的后续问题。在本文中,我们关注一个更受限的设置,即仅离线访问LLM,SLM参数保持固定,在线推理仅由SLM执行。为此,我们提出可重用潜在修正(RLC),将黑盒LLM的一次性自然语言指导转换为SLM隐藏空间中的持久修正经验。RLC将这些经验存储在外部库中,并根据SLM当前的推理状态检索它们,使SLM在推理过程中无需任何在线LLM调用即可重用LLM派生的修正。在多个推理基准和SLM规模上的实验表明,RLC在无需参数更新或在线LLM调用的情况下持续提高SLM推理能力。代码可在以下URL获取。

英文摘要

Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (SLMs) are easier to deploy locally yet remain weaker in reasoning. This capability-deployment gap has motivated LLM-SLM collaboration, which aims to improve SLM reasoning using LLM capabilities while preserving the deployment advantages of SLMs. Existing approaches mainly follow two paradigms. Knowledge distillation uses LLM-generated answers and reasoning trajectories to train SLMs offline, but requires parameter updates and additional training. Alternatively, online collaboration routes difficult problems to an LLM or leverages LLM-generated guidance and corrections when an SLM encounters difficulties. Although effective, online collaboration requires repeated LLM access. Moreover, the guidance produced for a particular problem is discarded after inference and cannot benefit subsequent problems involving similar reasoning states. In the paper, we focus on a more constrained setting in which the LLM is accessed only offline, the SLM parameters remain fixed, and online inference is performed solely by the SLM. To this end, we propose Reusable Latent Correction (RLC), which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM. RLC stores these experiences in an external bank and retrieves them according to the SLM's current reasoning state, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls. Experiments across multiple reasoning benchmarks and SLM scales show that RLC consistently improves SLM reasoning without parameter updates or online LLM calls. Code is available at https://github.com/ZBH031/reusable-latent-correction.

发表机构

  • Southeast University(东南大学)
  • Hong Kong Polytechnic University(香港理工大学)
  • ZTE Corporation(中兴通讯股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑