消费系统上投机解码中的执行路径限定与实现成本
Execution-Path Qualification and Realized Costs in Speculative Decoding on Consumer Systems
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究在消费级B570系统上评估投机解码的执行路径限定与成本,提出基于公开数据的预测器匹配72/72序列,但无法恢复限定,且成本分析显示RTX序列化较慢,为无损兼容性提供诊断。
AI中文摘要:
数值执行路径可以改变贪心输出;仅凭不匹配并不能预测剩余轨迹。在测试的B570堆栈上,我们利用公开的发现数据选择了一个无草稿、仅目标、微批次为一的预测器,然后在保留的实时投机之前冻结完整的生成ID预测。它在24个可重复稳定提示上匹配了72/72个封顶序列,包括四个参考发散提示的所有12次运行;同一集合的参考匹配为60/72。该预测器对此路径是充分的,但它既不能恢复限定,也不能确定唯一的数值原因。单独的B570同状态分叉测试干预效果;RTX提案/重放程序在共享数值例程的同时测试复现。我们分别评估成本。新近直接ID限定的RTX序列化在其仪器化的预填充后边界内较慢(目标/序列化延迟的$0.646\ imes$)。在固定的测量进度和目标工作下,即使移除草稿和残余成本,358H仍低于平价。合格的M4 Guard优于仅目标,但输给固定投机。这些结果为测量路径上的无损兼容性和完整成本提供了可测试的诊断和工程决策;它们未建立可移植的解码或控制增益。
英文摘要:
Numerical execution paths can change greedy output; a mismatch alone does not predict the remaining trajectory. On the tested B570 stack, we select a draft-free target-only micro-batch-one predictor using exposed discovery data, then freeze complete generated-ID predictions before reserved live speculation. It matches 72/72 capped sequences on 24 repeat-stable prompts, including all 12 runs of four reference-divergent prompts; the same-set reference matches 60/72. The predictor is sufficient for this path, but it neither restores qualification nor identifies a unique numerical cause. Separate B570 same-state forks test intervention effects; RTX proposal/replay programs test reproduction while sharing numerical routines. We assess cost separately. Newly direct-ID-qualified RTX serialization is slower within its instrumented post-prefill boundary ($0.646\times$ target/serialized latency). At fixed measured progress and target work, 358H remains below parity even with draft and residual costs removed. Qualified M4 Guard beats target-only but loses to fixed speculation. These results give testable diagnoses and engineering decisions about lossless compatibility and complete cost on the measured paths; they establish no portable decoding or control gain.