发表机构
SCBX R&D; SCB DataX, SCBX Group(SCBX研发中心; SCBX集团SCB DataX)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SpeechConversationBench(SCB)基准,用103个分片GSM8K问题评估语音到语音模型的多轮口语数学推理,比较完整、拼接和分片三种条件,发现分片导致商业系统准确率下降5.0-25.3个百分点,而LEGO在三种条件下均达77.5%准确率。
AI 中文摘要
语音到语音系统必须解决其需求在对话轮次中出现的任务。我们引入了SpeechConversationBench(SCB),这是一个使用103个分片的GSM8K问题对口语数学推理进行聚焦评估的基准。该框架比较了在一轮中传递的原始问题(完整)、其拼接的信息分片一起传递(拼接)以及跨轮次增量口语披露(分片)三种情况。我们报告了四个商业语音系统和LEGO(SCBX创新实验室团队内部开发的一种具有显式对话上下文管理的专有语音流水线)的最终答案准确率。相对于拼接,四个商业系统的分片准确率下降了5.0-25.3个百分点。LEGO在所有三种条件下均达到77.5%的准确率,而GPT-4o Realtime的分片准确率为76.6%。两个单轮基线区分了对问题重新表述的敏感性与增量口语交互带来的额外挑战。
英文摘要
Speech-to-speech systems must solve tasks whose requirements emerge across conversational turns. We introduce SpeechConversationBench (SCB), a focused evaluation of spoken mathematical reasoning using 103 sharded GSM8K problems. The framework compares the original problem delivered in one turn (full), its concatenated information shards delivered together (concat), and incremental spoken disclosure across turns (sharded). We report final-answer accuracy for four commercial speech systems and LEGO, a proprietary speech pipeline developed internally by the SCBX Innovation Lab team with explicit conversational context management. Relative to concat, sharded accuracy decreases by 5.0-25.3 percentage points across the four commercial systems. LEGO achieves 77.5 percent accuracy in all three conditions, compared with 76.6 percent sharded accuracy for GPT-4o Realtime. The two single-turn baselines distinguish sensitivity to problem reformulation from the additional challenges introduced by incremental spoken interaction.
CommentsConducted during a 2024 internship at SCBX R&D