发表机构
University at Buffalo; University of Texas at San Antonio(布法罗大学; 德克萨斯大学圣安东尼奥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OneVoice提出一种轻量级、原生JSON的中间表示,为智能体语音流水线提供共享语义结构,显著提升多智能体工作流中的聚合可靠性和时序关系保持。
AI 中文摘要
智能体语音系统必须交换超越文本的信息,包括说话人身份、时序、语音学信息和行为标注。然而,这些信号通常以不兼容的、特定于工具的形式产生,使得智能体之间的交接变得脆弱。我们提出OneVoice,一种轻量级、原生JSON的中间表示,为语音流水线提供共享的语义结构。OneVoice将异构语音证据组织成具有稳定标识符、分层转录、明确时序关系和来源信息的经过验证的会话记录。我们在两个互补的多智能体工作流中评估OneVoice,涵盖事件聚合和跨三种语言模型的声学-语音时序关联。与隐式的、智能体定义的交接相比,OneVoice显著提高了聚合可靠性和时序关系的保持,证明了显式的、语音特定的表示对智能体通信的价值。
英文摘要
Agentic speech systems must exchange more than text, including speaker identity, timing, phonetic information, and behavioral annotations. Yet these signals are often produced in incompatible tool-specific formats, making agent-to-agent handoff fragile. We present OneVoice, a lightweight, JSON-native intermediate representation that provides a shared semantic structure for speech pipelines. OneVoice organizes heterogeneous speech evidence into validated session records with stable identifiers, layered transcripts, explicit timing relationships, and provenance information. We evaluate OneVoice in two complementary multi-agent workflows covering event aggregation and acoustic-phonetic temporal linking across three language models. Compared with implicit agent-defined handoffs, OneVoice substantially improves aggregation reliability and the preservation of temporal relationships, demonstrating the value of an explicit speech-specific representation for agent communication.