发表机构
National Research Council Canada(加拿大国家研究委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PACT-SLM契约测试,用于评估流式口语智能体在部分语音前缀下采取行动的时机与身份,发现行动身份与时机衡量决策行为的不同维度。
AI 中文摘要
流式口语智能体可能在可用语音支持其行动之前就采取外部行动,然而最终轮次得分并不能揭示每个观察到的前缀是否支持该行动。我们引入了语音语言模型中轮流说话的局部语音行动契约(PACT-SLM),这是一种受控评估方法,它分配首个有效行动时间,并分别衡量行动身份和时机。主要诊断包含来自四个保留语义族的80个配对对比组,以及在干净和15分贝噪声渲染下的1,600个前缀预测。在纠正随机化分支代码与语义标签之间的不匹配后,重新拟合的WavLM Base Plus探针在发病后语义标签准确率上达到26.03%(组自助法95%置信区间:22.14%-29.68%),在发病前前缀上暴露了18.99%的行动,并精确预测了5.94%的完整轨迹。它在发病后标签准确率上超过了匹配的文本、标量声学和洗牌表示探针,但其得分位于100个前缀内标签排列的第96百分位,低于97.5百分位参考值(26.73%)。经过的时间比WavLM Base Plus更精确地对应发病时刻(36.25%对23.13%),但在行动身份上准确性较低(9.92%对26.03%)。这些结果表明,行动身份和时机衡量了局部语音决策行为的不同方面。
英文摘要
Streaming spoken agents may produce the correct final action after acting too early. Final-turn scores do not reveal whether each observed speech prefix supports an exposed action. We introduce the Partial Speech Action Contract for Turn Taking in Speech Language Models (PACT-SLM), a controlled test that assigns the first valid action time and evaluates both action identity and timing. In the primary test, 80 paired contrast groups from four held-out semantic families yield 1,600 prefix predictions across clean and 15 dB noise renderings. Using source-utterance semantic targets rather than counterbalanced branch codes, WavLM Base Plus reaches 26.03% pooled post-onset semantic-label accuracy (95% group-bootstrap interval [22.14%, 29.68%]), exposes an action on 18.99% of pre-onset prefixes, and predicts 5.94% of complete trajectories exactly. Its pooled label score is at the 96th percentile of 100 within-prefix label permutations, below the 97.5th-percentile reference (26.73%). It exceeds matched text, scalar-acoustic, and shuffled-representation probes in pooled post-onset label accuracy. Elapsed time is more onset-exact (36.25% versus 23.13%) but less accurate about action identity (9.92% versus 26.03%). These results motivate separate measurement of action identity and onset timing in partial-speech evaluations.