AI 中文总结
研究汽车应用中语音到语音大语言模型助手的保障措施,探讨基于转录本和工具的两种实现方法,经实证评估发现因高延迟和技术障碍,这两种策略大多不适用于工业部署,还概述了面临的开放挑战。
AI 中文摘要
近期进展引入了能够产生自然交互语音的语音到语音(S2S)对话助手,包括语调、情绪等非语言线索,在汽车领域能实现直观、类人的车内对话体验。但集成这些端到端助手限制了可编程领域特定保障措施的架构选择。本文讨论了S2S护栏的两种实现方法:基于转录本和基于工具的。通过实证评估,发现因高延迟(即使是计算成本低的检查,每个答案延迟0到1.4秒)和技术障碍(如潜在的非确定性工具调用行为),这两种策略在大多数情况下都不足以用于工业部署。最后概述了汽车环境中S2S护栏面临的开放挑战。
英文摘要
Recent advances have introduced speech-to-speech (S2S) conversational assistants capable of producing natural-sounding interactions, including non-verbal cues like tonality and mood. In the automotive domain, this enables intuitive and humanlike in-car dialogue experiences. However, integrating these end-to-end assistants limits architectural options for programmable domain-specific safeguards. This paper discusses two implementation approaches for S2S guardrails: transcript-based and tool-based. Through an empirical evaluation, we demonstrate that both strategies are insufficient for industrial deployment in most cases due to prohibitive latency (delaying each answer by 0 to 1.4 seconds even for computationally cheap checks) and technical impediments (like potentially non-deterministic tool call behavior). Finally, we outline open challenges for S2S guardrails in the automotive context.
Journal refLNCS Volume 16830, 2026
DOI:10.1007/978-3-032-32335-4_18