arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过推测执行隐藏端侧级联语音代理中的工具延迟

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park, Jinhyeok Yang, KiHyun Nam, Jaegul Choo, Jinkyu Lee

arXiv 2610.07641首次发表:更新:

发表机构

Qualcomm Technologies, Inc. (QTI)(高通技术公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对端侧级联语音代理,提出推测性工具执行方法,通过预测工具调用并缓存结果来隐藏工具延迟,实测中位首次音频响应时间从5.79秒降至4.60秒,延迟更稳定。

AI 中文摘要

工具增强型语音助手通常将自动语音识别、大语言模型推理和外部工具执行串行化。因此,工具延迟仅在用户说完话且大语言模型识别出所需工具调用之后才产生。我们提出了面向端侧级联语音代理的推测性工具执行方法,该方法从部分ASR假设中预测工具请求,并在语音仍在接收时启动工具执行,从而减少端到端响应延迟。我们的方法引入了一个预测器模块,在语音识别过程中预测工具调用,推测性地执行这些调用并缓存结果。随后将缓存输出注入大语言模型提示中,实现更快的响应。此外,为减轻用户说话时自我纠正引起的错误,我们采用基于规则的验证机制,仅选择性地注入有效的缓存结果。作为最终保障,大语言模型保留直接发出工具调用的能力,确保在最坏情况下我们框架的延迟以基线串行执行流程为上限。我们使用完全实现的Android语音助手的实时测量来评估我们的方法。该方法将中位首次音频响应时间从5.79秒缩短至4.60秒,并将标准差从3.49秒降至2.81秒,从而使响应延迟更具可预测性。

英文摘要

Tool-augmented speech assistants typically serialize automatic speech recognition, large language model inference, and external tool execution. As a result, tool latency is incurred only after the user has finished speaking and the LLM has identified the required tool calls. We present speculative tool execution for on-device cascaded voice agents, which predicts tool requests from partial ASR hypotheses and initiates tool execution while speech is still being received, thereby reducing end-to-end response latency. Our approach introduces a Predictor module that anticipates tool calls during speech recognition, executes them speculatively, and caches the results. The cached outputs are then injected into the LLM prompt, enabling faster responses. Additionally, to mitigate errors caused by user self-corrections during speech, we employ a rule-based validation mechanism that selectively injects only valid cached results. As a final safeguard, the LLM retains the ability to issue tool calls directly, ensuring that the latency of our framework is upper-bounded by the baseline serial execution pipeline in the worst case. We evaluate our method using live measurements from a fully implemented Android voice assistant. Our approach reduces the median time-to-first-audio from 5.79,s to 4.60,s and decreases the standard deviation from 3.49,s to 2.81,s, resulting in more predictable response latency.

Comments5 pages, 2 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑