arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将CPU语音合成推向极限:无服务器架构与计费下的极端推理调优

Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing

Pakorn Nathong, Kunat Pipatanakul

arXiv 2610.00063首次发表:更新:

发表机构

Paxa Labs(帕萨实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无服务器按实例计费下CPU推理空闲成本高的问题,提出计费感知的TTS服务,通过请求级并行和可回收生命周期,在Kokoro-82M上实现4.1倍成本降低。

AI 中文摘要

按实例计费的无服务器平台在热实例的整个生命周期内对CPU和内存收费,这使得空闲推理状态成为直接的运维成本。我们提出了在无服务器CPU上进行计费感知的神经文本转语音(TTS)服务,优化CPU秒和GB秒而非单纯的吞吐量或延迟。传统运行时不适合此场景:每请求并行度在并发下导致CPU争用,而热实例保留数GB的计费推理状态和页缓存。我们通过请求大小的并发推理(限制每请求CPU并行度)和可回收的实例生命周期(在空闲期后释放推理状态和页缓存内存,同时保留服务器进程和编译缓存)来解决这些成本。在Kokoro-82M上,我们的系统达到每CPU秒2.71音频秒,而ONNX Runtime默认仅为0.89,并将每音频小时成本从PyTorch的0.0631美元降至0.0153美元,降低4.1倍。空闲计费内存从8.7 GB降至1.33 GB,恢复时首次音频输出仅需2.2秒,而PyTorch冷启动为7.7秒。在突发流量下,生命周期回收对于将推理效率转化为更低的无服务器成本至关重要。

英文摘要

Instance-billed serverless platforms charge for CPU and memory over the lifetime of a warm instance, making idle inference state a direct serving cost. We present billing-aware neural text-to-speech (TTS) serving on serverless CPUs, optimizing CPU-seconds and GB-seconds rather than throughput or latency alone. Conventional runtimes are poorly suited to this setting: per-request parallelism causes CPU contention under concurrency, while warm instances retain gigabytes of billable inference and page-cache state. We address these costs with request-sized concurrent inference, which bounds per-request CPU parallelism, and a reclaimable instance lifecycle, which releases inference state and page-cache memory after idle periods while retaining the server process and compile cache. On Kokoro-82M, our system achieves 2.71 audio-seconds per CPU-second versus 0.89 with ONNX Runtime defaults and reduces cost per audio-hour from $0.0631 with PyTorch to $0.0153, a 4.1x reduction. Idle billed memory falls from 8.7 GB to 1.33 GB, while restoration reaches first audio in 2.2 s versus 7.7 s for a PyTorch cold start. Under bursty traffic, lifecycle reclamation is essential for translating inference efficiency into lower serverless cost.

Comments6 pages, technical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑