PyroDash:具有成本效益的令牌级小语言模型与大语言模型协作推理
PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference
- Pyromind Dynamics Inc.(Pyromind动力学公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对大语言模型推理成本高、小语言模型可靠性低的问题,提出PyroDash框架,通过令牌级协作推理,分三阶段训练小语言模型,在数学推理基准测试中能支持不同操作点,可减少大语言模型使用并保持推理性能。
AI中文摘要:
大语言模型(LLMs)推理能力强但大规模服务成本高,小语言模型(SLMs)成本低但处理难题可靠性差。我们引入了PyroDash,一个用于令牌级SLM-LLM协作推理的成本感知框架。生成过程中,SLM通过发出控制令牌决定是否请求协助,协作引擎将查询和部分推理轨迹发送给冻结的LLM完成。PyroDash分三个阶段训练SLM,其奖励平衡答案准确性和仅使用LLM推理归一化后的推理成本。在五个数学推理基准测试中,PyroDash支持不同的准确性-成本操作点,结果表明学习到的令牌级交接可减少LLM使用并保持强大推理性能。
英文摘要:
Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware framework for token-level SLM-LLM collaborative inference. During generation, the SLM decides whether to request assistance by emitting a control token. A Collaborate Engine then sends the query and partial reasoning trace to a frozen LLM for completion through a single handoff. The policy is internalized in the SLM, requiring neither a separate router, LLM retraining, nor access to LLM logits. PyroDash trains the SLM in three stages: control-token embedding learning, offloading-oriented supervised fine-tuning, and cost-aware alignment with Group Relative Policy Optimization. Its reward balances answer accuracy against inference cost normalized by LLM-only inference. Across five mathematical reasoning benchmarks, PyroDash supports different accuracy-cost operating points. With $λ=0.05$, it achieves 64.04 percent average accuracy, 6.36 percentage points above the LLM-only baseline, while reducing cost by 20.4 percent. With $λ=0.6$, it achieves 54.55 percent accuracy with a 1.90 percent LLM token ratio and 0.012 LLM calls per example, reducing total cost from USD 49.36 to USD 1.78. These results show that learned token-level handoffs can reduce LLM use while preserving strong reasoning performance.