发表机构
Zuoyebang Education Technology(作业帮教育科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全双工对话中话轮转换评估忽视上下文的问题,提出配对基准ECHO和配对准确率指标,揭示多数系统偏向让出,打断评估会高估可靠性。
AI 中文摘要
全双工口语对话系统必须区分需要让出话语权的打断和允许继续说话的反馈信号。现有基准通常独立评估事件,因此可能奖励固定的动作偏好而非上下文敏感的决策。我们引入了ECHO,一个用于中文全双工话轮转换的配对诊断基准。ECHO将具有相同重叠转录但对比性多轮对话上下文的示例配对,其中一个需要让出(Yield),另一个需要保持(Keep)。它还包括用于诊断不必要让出的离题(off-talk)示例。我们引入了配对准确率,该指标要求在一对示例的两个成员上都做出正确决策,并且对恒定动作策略不给予任何分数。对多个全双工系统的实验表明,大多数系统表现出明显的让出(Yield)偏向,在打断上的表现显著优于在反馈信号上的表现,而另一个系统则相对平衡。这些发现表明,仅基于打断的评估可能高估实际的话轮转换可靠性。ECHO及其元数据将公开发布。
英文摘要
Real-time spoken dialogue systems must distinguish interruptions that require yielding the floor from backchannels that permit continued speaking. Existing benchmarks typically score events independently and may therefore assign high scores to systems with fixed action preferences rather than context-sensitive decision policies. We introduce ECHO, a paired diagnostic benchmark for Chinese turn-taking evaluation. ECHO pairs examples with the same overlap transcript but contrasting preceding multi-turn dialogue contexts, with one requiring Yield and the other Keep. It additionally includes off-talk examples for diagnosing unnecessary yielding. We introduce pair accuracy, which requires correct decisions on both members of a pair and assigns no credit to constant-action policies. Experiments on four speech systems show that three exhibit a severe over-yielding bias: they correctly keep the floor on fewer than 13% of backchannels, resulting in near-zero pairwise success rates equal or less than 4%. While the remaining system remains comparatively balanced across contexts, these findings broadly demonstrate that interruption-only evaluation can severely overestimate practical turn-taking reliability.