arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向语音的LLM服务的可扩展上下文编排

Scalable Context Orchestration for Serving LLMs Over Voice

Linyi Jiang, Silvery D. Fu, Yifei Zhu

arXiv 2609.04288首次发表:更新:

发表机构

Global College, Shanghai Jiao Tong University; AgenticSys(上海交通大学Global College; AgenticSys)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出llmovoice中间件,通过明确建模和编排语音上下文,降低语音LLM服务的语速对齐误差、误中断率和成本,同时保留高答案质量。

AI 中文摘要

随着大语言模型(LLM)的进步使语音交互更自然、更易获取,语音AI应用日益普及。服务这些应用不仅需要考虑用户所说的内容,还需考虑用户的说话方式(如语速)以及音频捕获和传输的条件(如背景噪声和丢包)。然而,现有的LLM系统将对话上下文表示为不断增长的扁平消息序列,将语音特定上下文隐含在音频中,导致生成的响应与用户偏好匹配度差、在不利环境条件下交互质量下降、长时间语音会话成本高昂。我们提出llmovoice,一种明确建模语音上下文并编排其使用的上下文管理中间件。每一轮,llmovoice从当前用户输入、相关交互历史以及明确的副语言和环境状态构建有界语音上下文,然后利用服务LLM对该上下文进行推理,生成指导系统响应方式的运行时指令。我们在真实世界语音应用和基准上评估llmovoice,结果显示:它将语速对齐误差降低52.4%,在丢包情况下将误中断率从46.0%降至0.9%,模型使用成本降低79.2%;对于长时间会话,llmovoice在保留多达98.7%基线答案质量的同时,每轮成本最多降低24.9倍。

英文摘要

Voice AI applications are gaining popularity as advances in large language models (LLMs) enable more natural and accessible spoken interactions. Serving these applications requires accounting not only for what users say, but also for how they speak (e.g., speaking rate) and the conditions under which their audio is captured and transmitted (e.g., background noise and packet loss). However, existing LLM systems represent conversation context as a flat, growing sequence of messages, leaving voice-specific context implicit in the audio. As a result, they can generate responses that are poorly aligned with user preferences, degrade interaction quality under adverse environmental conditions, and incur high costs over long voice sessions. We present llmovoice, a context-management middleware that explicitly models voice context and orchestrates its use. At each turn, llmovoice constructs a bounded voice context from the current user input, relevant interaction history, and explicit paralinguistic and environmental states. It then uses the serving LLM to reason over this context and generate runtime directives that guide how the system responds. We evaluate llmovoice on real-world voice applications and benchmarks. It reduces speaking-rate alignment error by 52.4%, lowers the false-interruption rate from 46.0% to 0.9% under packet loss, and reduces model usage cost by 79.2%. For long sessions, llmovoice reduces per-turn cost by up to 24.9 times while retaining up to 98.7% of baseline answer quality.

CommentsAccepted for publication in ACM SOSP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑