发表机构
Jinan University; Hong Kong University of Science and Technology (Guangzhou)(暨南大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DuplexCadence通过向运行时声明语音模型的本地时钟,实现按需状态分配和精确形状图重放,解决全双工语音服务中的编排空闲与内存浪费,速度提升2.85倍,峰值内存降低38.8%。
AI 中文摘要
全双工语音模型支持同时进行听和说的流式交互。其服务受到严格且重复的截止时间约束:对话以一秒为节奏推进,每一秒的输入必须在下一秒到来之前转化为一秒的语音。由于会话内的各阶段严格顺序执行,每次调用的开销无法通过批处理消除。性能分析显示,双工一秒内的自回归阶段已能适应周期,而令牌到音频的合成尾部是导致超时的原因。该尾部阶段存在编排空闲问题,GPU在逐条发出数千个微小且规律的操作时处于等待状态,同时因在静态实现常量上过度配置状态而浪费大量内存。现有补救措施(如图记录和按需大小分配)因流式状态动态违反其前提而失效。根本原因在于运行时缺乏模型的本地时钟:即控制推进速率和保留策略的每区域计数器。我们提出DuplexCadence,它向运行时显式声明本地时钟,并推导出两个相互促进的规则:在稳定地址进行按需大小的状态分配,以及无填充的精确形状图重放。前者消除空闲内存并稳定张量指针,后者在无填充开销的情况下消除编排空闲。在四个已发布模型、三种解码器架构上评估,且输出逐位相同,DuplexCadence达到标准运行时2.85倍的速度,峰值内存降低38.8%。在实时双工路径上,平均SPEAK时间从超过一秒节奏的14%降至低于该节奏的2%,使模型能够可靠地跟上交互式语音,同时显著扩展多任务能力。
英文摘要
Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. Because stages within a session run in strict sequence, per-invocation overhead cannot be batched away. Profiling reveals that the autoregressive stages of a duplex second already fit within the period, whereas the token-to-audio synthesis tail is what causes overruns. This tail stage suffers from orchestration slack where the GPU is left waiting as thousands of tiny, regular operations are issued one by one, while also wasting substantial memory by over-provisioning state at static implementation constants. Existing remedies, such as graph recording and demand-sized allocation, fail because streaming state dynamics violate their prerequisites. The root cause is that the runtime lacks the model's native clocks: the per-region counters that govern advancement rates and retention policies. We propose DuplexCadence, which explicitly declares native clocks to the runtime and derives two mutually enabling rules: demand-sized state allocation at a stable address, and exact-shape graph replay without padding. The former eliminates idle memory and stabilizes tensor pointers, while the latter removes orchestration slack without padding overhead. Evaluated on four released models across three decoder architectures with bit-for-bit identical output, DuplexCadence reaches $2.85\times$ the stock runtime's speed at $38.8\%$ lower peak memory. On the live duplex path, mean SPEAK time falls from $14\%$ over the one-second cadence to $2\%$ under it, enabling models to reliably keep up with interactive speech while markedly expanding multi-