arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学会提问:预算约束下的SLM-LLM协作信息获取

Learning to Ask: Information Acquisition for SLM-LLM Collaboration, under a budget

Yongjun Kim, Xiaoxiao Li, Jaeho Lee

arXiv 2610.01236首次发表:更新:

发表机构

Pohang University of Science and Technology (POSTECH); University of British Columbia; Vector Institute(浦项科技大学; 不列颠哥伦比亚大学; 向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将SLM-LLM协作视为预算约束下的信息获取问题,提出三阶段RLVR框架,使SLM选择性查询LLM顾问,在数学推理和编码任务中优化性能-成本权衡,并可迁移至其他模型。

AI 中文摘要

小语言模型(SLM)与大语言模型(LLM)之间的协作提供了一个机会,可以将较小模型的效率与较大模型的强大推理能力相结合。现有方法主要将这种协作视为计算分配问题,即确定哪个模型应处理推理过程的每个部分。然而,在基于黑盒API的设置中,由于粗粒度的委派或模型切换时上下文的重复传输,这种范式可能效率低下。在这项工作中,我们转而将SLM-LLM协作表述为在API预算约束下的信息获取问题。SLM保持为主要推理者,仅在需要时选择性地查询黑盒LLM顾问,发出有针对性的查询,而不是委派推理过程本身。为实现这一策略,我们开发了一个三阶段RLVR框架,通过联合优化顾问调用和信息使用,学习是否调用顾问、如何制定有用的查询以及如何将协作整合到推理过程中。在数学推理和编码任务中,我们的方法在性能-成本权衡上优于现有的协作基线,并且在某些设置中,达到或超过了oracle问题级路由的性能。最后,我们展示了我们的策略可以无需进一步训练即可迁移到其他顾问模型家族。

英文摘要

Collaboration between a small language model (SLM) and a large language model (LLM) offers an opportunity to combine the efficiency of smaller models with the strong reasoning capabilities of larger ones. Existing approaches primarily frame such collaboration as a computation allocation problem, determining which model should handle each portion of the reasoning process. In black-box API-based settings, however, this paradigm can be inefficient due to coarse-grained delegation or repeated transmission of context across model switches. In this work, we instead formulate SLM-LLM collaboration as an information acquisition problem, under an API budget constraint. The SLM remains the primary reasoner and selectively queries a black-box LLM advisor only when needed, issuing targeted queries rather than delegating the reasoning process itself. To realize this strategy, we develop a three-stage RLVR framework that learns whether to call the advisor, how to formulate useful queries, and how to integrate the collaboration into the reasoning process by jointly refining advisor invocation and information use. Across mathematical reasoning and coding tasks, our approach improves the performance--cost tradeoff over existing collaboration baselines and, in some settings, matches or exceeds oracle problem-level routing. Finally, we show that our strategy can transfer to other advisor model families, without further training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑