$S^3$-Bench:评估作为科学语音助手的语音交互模型
$S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants
- Shanghai Jiao Tong University(上海交通大学)
- Ant Group(蚂蚁集团)
- Tongji University(同济大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出S$^3$-Bench,一个覆盖10个学科的系统评估框架,通过知识集和对话集分解语音交互阶段,揭示科学语音助手在术语处理、多轮适应及响应质量上的挑战与权衡。
AI中文摘要:
多模态大语言模型(MLLMs)的进步从根本上重塑了人机交互的范式,尤其是能够进行无缝对话的语音交互模型。尽管作为通用语音助手表现出色,但它们在专业领域(尤其是科学领域)的性能仍未得到充分探索。科学交互带来了严峻的挑战,涉及罕见的技术术语、缩写的口语规范以及符号特殊表达的自然言语化。在本文中,我们引入了S$^3$-Bench,一个覆盖10个主要学科的系统性评估框架,包含用于语音问答的知识集和用于与模拟用户代理进行多轮渐进式交互的对话集。通过将完整的原子轮次分解为语音识别、感知、知识利用与推理以及响应发音等阶段,我们系统地刻画了现有方法的常见挑战和性能权衡。此外,多轮交互实验揭示了在用户适应以及生成准确、全面且高效响应方面的持续局限性。
英文摘要:
The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their performance in specialized domains remains underexplored, particularly in scientific areas. Scientific interactions introduce formidable challenges, involving rare technical terminology, spoken norms of abbreviations, and the natural verbalization of symbolic special expressions. In this paper, we introduce S$^3$-Bench, a systematic evaluation framework covering 10 major disciplines, consisting of a Knowledge set for speech question-answering and a Dialogue set for multi-turn progressive interactions with simulated user agents. By decomposing a complete atomic turn into stages of speech recognition, perception, knowledge utilization with reasoning, and response pronunciation, we systematically characterize the common challenges and performance tradeoffs of existing approaches. Furthermore, experiments on multi-turn interactions reveal persistent limitations in user adaptation and the generation of accurate, comprehensive, and efficient responses.