arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AS$^2$D:加速移动设备上的按需音频理解

AS$^2$D: Accelerating On-Demand Audio Understanding on Mobile Devices

Yunzhe Li, Kyoungjun Park, Hongzi Zhu, Lili Qiu

arXiv 2609.37617首次发表:更新:

发表机构

The University of Texas at Austin; Shanghai Jiao Tong University(德克萨斯大学奥斯汀分校; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对移动设备音频理解,提出AS$^2$D解码方法,使草稿生成与目标验证解耦并行,在四款手机上提升ASR吞吐量42-76%,显著优于投机基线。

AI 中文摘要

投机解码通过使用较小的草稿模型为较大的目标模型提出令牌,以供批量验证,从而加速自回归生成。然而,传统的投机解码将草稿生成与目标模型不断演化的已验证前缀耦合在一起,使草稿生成和验证串行化。我们质疑这种依赖关系对于源条件生成是否必要。我们的关键观察是,对于音频语言模型,输入音频和用户请求可以提供有用的投机候选,而无需遵循目标模型不断演化的文本前缀。我们提出了AS$^2$D(音频投机投机解码),它实现了目标解耦的草稿生成:一个音频条件的草稿模型遵循其自身的生成历史,而目标模型独立验证并纠正准备好的候选。在没有可用候选时,目标模型单独推进。因此,目标反馈决定哪些候选被提交,但不再决定草稿模型何时能取得进展,从而使草稿生成和验证能够并发进行,同时保留目标侧的验证和纠正。我们在MNN中为Android实现了AS$^2$D,并在四款手机、七个数据集和三个任务上评估了两个目标模型,覆盖了12.2小时的音频。在四款手机上,AS$^2$D相比仅目标解码将池化ASR吞吐量提高了42-76%,而只有5.7%的评估窗口比仅目标解码慢,相比之下投机基线为58.1-63.0%。对于ASR,AS$^2$D在评估的草稿模型/预算目录中达到了事后逐窗口预言机池化吞吐量的97.33-98.20%。使用7B目标模型的原生按需执行比仅目标解码的吞吐量高出最多78%。这些结果表明,源条件的音频生成可以放宽投机草稿生成对目标模型不断演化的输出前缀的传统依赖,从而为高效推理暴露大量并行性。

英文摘要

Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency is necessary for source-conditioned generation. Our key observation is that, for audio language models, the input audio and user request can provide useful speculative candidates without following the target's evolving text prefix. We propose AS$^2$D (Audio Speculative Speculative Decoding), which enables target-decoupled drafting: an audio-conditioned drafter follows its own generation history while the target independently verifies and corrects ready candidates. Without usable candidates, the target advances alone. Thus, target feedback determines which candidates are committed but no longer determines when the drafter can make progress, enabling drafting and verification to proceed concurrently while retaining target-side verification and correction. We implement AS$^2$D in MNN for Android and evaluate two target models across four phones, seven datasets, and three tasks covering 12.2 hours of audio. Across four phones, AS$^2$D improves pooled ASR throughput by 42-76% over target-only decoding, while only 5.7% of evaluation windows are slower than target-only, compared with 58.1-63.0% for speculative baselines. For ASR, AS$^2$D reaches 97.33-98.20% of a hindsight per-window oracle's pooled throughput over the evaluated drafter/budget catalog. Native on-demand execution with a 7B target achieves up to 78% higher throughput than target-only. These results show that source-conditioned audio generation can relax the conventional dependence of speculative drafting on the target's evolving output prefix, exposing substantial parallelism for efficient inference.

Comments43 pages, 9 figures, 16 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑