arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FutureBridge:协作解码中超越局部偏好的令牌选择

FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding

Quanquan Li, Hongbo Zhang, Yihe Chi, Jingyu Li, Xidong Xi, Liuyang Song, Hongzhen Zhang, Yuxiang Huang, Jing Ke, Siyuan Ma, Junyi Lin, Guitao Cao

arXiv 2608.06819首次发表:更新:

AI 中文总结

FutureBridge通过联合LLM-SLM令牌对SLM后续推理的支持度排序,在5个数学推理基准上使Qwen3-1.7B SLM的数学平均得分较贪心解码提升35.1%,超越了仅依赖LLM局部偏好的现有方法。

AI 中文摘要

令牌级协作允许大型语言模型(LLM)在与小型语言模型(SLM)预测不一致时为SLM提供辅助。现有方法要么使用LLM生成的干预令牌,要么通过LLM的下一个令牌概率对候选进行排序,二者均依赖LLM的局部偏好,尽管LLM所选令牌可能难以被SLM用于后续推理。本文提出FutureBridge,该方法根据LLM与SLM的联合令牌候选对SLM后续推理的支持程度对其排序。训练期间,经答案验证的LLM轨迹提供固定的共享未来,冻结的SLM在该共同上下文下评估每个候选,所得反事实分数监督仅观察当前状态和候选令牌的轻量级令牌重排序器。推理时,FutureBridge仅用LLM扩展候选池,选择一个令牌后将生成权返回给SLM,不生成或附加未来后缀。在五个数学推理基准测试中,FutureBridge使Qwen3-1.7B SLM的数学平均得分较贪心SLM解码提升35.1%。这些结果表明,令牌选择需建模接收方SLM能否利用每个候选继续推理,而非仅依赖LLM的局部偏好。

英文摘要

Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existing methods either use LLM-generated intervention tokens or rank candidates with the LLM's next-token probabilities. Both rely on the LLM's local preference, even though an LLM-selected token may be difficult for the SLM to build on. We present FutureBridge, which ranks joint LLM-SLM token candidates according to how well they support the SLM's subsequent reasoning. During training, an answer-verified LLM trajectory supplies a fixed shared future, and a frozen SLM evaluates every candidate under this common context. The resulting counterfactual scores supervise a lightweight token reranker that observes only the current state and candidate token. At inference, FutureBridge uses the LLM only to expand the candidate pool, selects one token, and returns generation to the SLM without generating or appending a future suffix. Across five mathematical reasoning benchmarks, FutureBridge improves the Qwen3-1.7B SLM's Math Avg. by 35.1% relative to greedy SLM decoding. These results indicate that token selection benefits from modeling whether the receiving SLM can use each candidate to continue reasoning, rather than relying on the LLM's local preference alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑