arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TSS:面向领域特定大语言模型的推测解码目标端稀疏化

TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

Haibo Hu, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue

arXiv 2609.26100首次发表:更新:

发表机构

City University of Hong Kong; Mohamed bin Zayed University of Artificial Intelligence(香港城市大学; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对领域特定大语言模型推测解码,提出目标端稀疏化框架TSS,通过跳过选定层降低验证成本并提升接受率,在翻译任务中实现1.68倍加速。

AI 中文摘要

推测解码通过轻量级草稿模型与目标验证器之间的协作来加速大语言模型推理。现有方法主要改进草稿端,而目标模型通常保持稠密且不变。我们表明,在领域特定推理下,全深度目标验证并不总是最优选择。与直觉相反,跳过选定的目标层可以降低验证成本,同时提高草稿接受率,并保持甚至提升下游任务性能。基于这一观察,我们提出TSS,一种用于推测解码的目标端稀疏化框架。TSS采用接受率和指标感知的广度搜索来探索多层跳过配置,而不在两个目标之间施加固定优先级。选定的配置存储在领域到配置的映射中,并由轻量级跳过控制器应用,使得一个完整的目标模型能够支持多个稀疏验证路径,而无需重新训练或永久参数剪枝。在Spec-Bench上跨多个领域、模型规模和推测解码方法的实验显示,草稿接受率和下游任务性能持续提升。在翻译设置中,TSS将平均接受长度从2.70提高到4.53(+67.8%),BLEU从0.131提高到0.237(+80.9%),端到端吞吐量从75.6提高到127.3 tokens/s,对应1.68倍加速。

英文摘要

Speculative decoding accelerates large language model inference through collaboration between a lightweight draft model and a target verifier. Existing methods mainly improve the draft side, while the target model is typically kept dense and unchanged. We show that, under domain-specific inference, full-depth target verification is not always the optimal choice. Counter-intuitively, skipping selected target layers can reduce verification cost while simultaneously increasing draft acceptance and preserving, or even improving, downstream task performance. Based on this observation, we propose TSS, a target-side sparsification framework for speculative decoding. TSS employs an acceptance- and metric-aware breadth search to explore multi-layer skip configurations without imposing a fixed priority between the two objectives. The selected configurations are stored in a domain-to-configuration mapping and applied by a lightweight skip controller, allowing one complete target model to support multiple sparse verification paths without retraining or permanent parameter pruning. Experiments on Spec-Bench across multiple domains, model scales, and speculative decoding methods show consistent improvements in draft acceptance and downstream task performance. In Translation setting, TSS increases the average accept length from 2.70 to 4.53 (+67.8%), improves BLEU from 0.131 to 0.237 (+80.9%), and raises end-to-end throughput from 75.6 to 127.3 tokens/s, corresponding to a 1.68X speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑