更快的块扩散推理服务,具有无分布风险保证
Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees
浏览论文内容
中文总结 AI 辅助
提出Redline,一种有限样本程序,在给定风险预算下选择最快且满足参考相对风险保证的块扩散模型操作点,实现更快推理并保持风险控制。
中文摘要 AI 辅助
块扩散语言模型在手工选择的操作点(如接受阈值、缓冲区深度、调度、检查点和精度)上提供服务,每个点都根据其平均基准准确率来选择。然而,平均值并不能告诉操作员,较快的配置在较慢配置能正确回答的提示上,多久会失败一次。在服务引擎及其解码轨迹上,默认提交规则已经提交了每个完全解析的块,静态跳过规则几乎捕获了分配所能节省的所有计算,而在引擎解码目标上的自蒸馏在准确率不变的情况下增加了速度。更大的加速来自更低的阈值,这会提交仍然不确定的令牌。因此,我们提出了Redline,一种有限样本程序,它根据校准提示上答案的正确性来选择操作点(手工选择或学习)。Redline将参考相对风险(即参考配置正确回答而候选配置不回答的联合概率)保持在用户选择的预算内,并以高概率部署通过检查的最快配置。在两种模型家族中,它在数学任务上以比代码任务更小的风险预算实现了加速,并且在百分之十的预算下,它部署了一个LLaDA2数学配置,该配置在每次前向传播中提交了超过三分之一的额外令牌。它还可以不加修改地应用于投机解码的接受规则和权重量化。在相同的校准数据上,Redline保持在其声明的失败概率内,而平均准确率规则的每个容差要么在某个模型和任务上获得更少的加速,要么在另一个模型和任务上更频繁地超出风险预算。代码可在以下网址获取:此https URL。
英文摘要
Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on prompts that the slower one answers correctly. On the serving engine and its decode traces, the default commit rule already commits every fully resolved block, a static skip rule captures nearly all of the compute that allocation can save, and self-distillation on engine-decoded targets adds speed at unchanged accuracy. Larger speedups come from lower thresholds, which commit tokens that are still uncertain. We therefore present Redline, a finite-sample procedure that selects operating points, hand-picked or learned, from the correctness of their answers on calibration prompts. Redline keeps the reference-relative risk, the joint probability that the reference answers correctly and a candidate configuration does not, within a user-chosen budget with high probability, and deploys the fastest configuration that passes. It speeds up math at a smaller risk budget than code in both model families, and at a budget of ten percent it deploys a LLaDA2 math configuration that commits over a third more tokens in each forward. It also applies without modification to the acceptance rule of speculative decoding and to weight quantization. On the same calibration data, Redline stays within its stated failure probability, whereas each tolerance of a mean-accuracy rule either gains less speed for some model and task or exceeds the risk budget far more often for another. Code is available at https://github.com/js-lee-AI/Redline.
发表机构
- Korea University(高丽大学)
- Zoom Communications(Zoom通信公司)
- Soongsil University(崇实大学)
- Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。