arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

既不沉默也不重叠是失败:全双工口语对话模型中基于意图的轮流说话评估

Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models

Kian Shamsaie, Iman Modarressi

arXiv 2609.27372首次发表:更新:

发表机构

People Make Things(People Make Things)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对全双工口语对话模型,提出意图条件评估基准TACT,用连续概率分数替代二元规则,在十一个系统中最佳模型达0.47,与人类判断高度一致。

AI 中文摘要

全双工口语对话模型的基准测试使用二元固定窗口规则来评分轮流说话,这些规则根据先前回合的完整性来奖励即时回应或沉默。我们认为,回应偏移的适当性——无论是延迟的沉默还是预期的重叠——取决于说话者的潜在意图,而这只能从该说话者的行为中识别。我们引入了TACT,一个包含来自五个双人语料库的9,728个片段和73.2小时数据的基准;每个片段都带有对话历史、每个说话者的记忆概况,以及注释者推导的六类意图后验分布。评分用严格适当的阈值加权连续排名概率分数取代了二元窗口,其权重是拟合人类话轮转换偏移分布的意图条件时间核,证明了有界性、一致性和二元约简性。在十一个系统中,最佳模型达到了0.47,而人类上限为0.86,且几乎不受说话者概况的影响,TACT与人类判断的一致性为Spearman 0.81,而二元指标为0.46。

英文摘要

Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.

CommentsAccepted to IEEE SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑