arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38386cs.AI

解码延迟反馈预填充:一种无模型控制器及其泛化极限

Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits

Gaurav Agarwal, Ashish Garg, Isha Singhal

AI总结:

针对并发推理中预填充干扰解码的问题,提出无模型控制器DLFP,通过比例反馈调整预填充块大小,在Qwen3-0.6B上降低P99延迟27.7%,但无法泛化至更大模型或并行配置。

AI中文摘要:

并发自回归推理会产生一个根本性的干扰问题:预填充新到达的长提示可能会延迟已在解码的请求的令牌输出。固定的预填充块可以减少这种干扰,但最佳块大小取决于模型、硬件、负载和延迟目标。我们引入了解码延迟反馈预填充(DLFP),这是一种无模型控制器,仅改变与活动解码重叠的预填充工作。在一个受保护的调度周期后,DLFP使用观察到的间隔作为比例反馈来调整下一个预填充块的大小;孤立的预填充保持不受限制。我们在vLLM中实现了DLFP,并使用开环泊松到达、精确令牌计数、原始请求轨迹和NVIDIA遥测数据对其进行了评估。在单块A100 80 GB GPU上使用BF16精度的Qwen3-0.6B模型,三次配对的100请求试验将P99令牌间延迟分别降低了24.8%、30.1%和28.2%(平均值27.7%,配对95%置信区间为21.0%至34.3%),且输出完全一致,无失败,SLO合规性不变。这种收益并非免费:平均P99首令牌时间增加了34.8%,但仍保持在声明的SLO之内。关键的是,该机制无法泛化到Qwen3-8B、Qwen3-32B或双GPU张量并行配置。我们将失败归因于异步调度器调用间隔,该间隔仅是完成的GPU迭代时间的代理。这一负面结果界定了贡献的边界,并激励了针对并发CPU和端侧设备推理的完成时间控制器。我们不声称移动设备性能;当前工作是一个可复现的概念验证和泛化研究。

英文摘要:

Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After a guarded scheduling cycle, DLFP uses the observed interval as proportional feedback to resize the next prefill chunk; isolated prefills remain unrestricted. We implement DLFP in vLLM and evaluate it with open-loop Poisson arrivals, exact token accounting, raw request traces, and NVIDIA telemetry. On Qwen3-0.6B in BF16 on one A100 80 GB GPU, three paired 100-request trials reduce P99 inter-token latency by 24.8%, 30.1%, and 28.2% (mean 27.7%, paired 95% confidence interval 21.0% to 34.3%) with exact output agreement, no failures, and unchanged SLO compliance. The benefit is not free: mean P99 time to first token increases 34.8% while remaining inside the declared SLO. Crucially, the mechanism does not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel configuration. We trace the failure to an asynchronous scheduler-call interval that is only a proxy for completed GPU iteration time. This negative result defines the boundary of the contribution and motivates a completion-timed controller for concurrent CPU and on-device inference. We do not claim mobile-device performance; the present work is a reproducible proof-of-concept and generalization study.

补充信息

↑