arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

区分语音语言模型中的决策规则失配与读出覆盖限制

Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

Linkai Peng, Baorian Nuchged

arXiv 2608.06409首次发表:更新:

发表机构

Institute for the Brain and Cognitive Sciences, University of Connecticut; Department of Linguistics, The University of Texas at Austin(康涅狄格大学脑与认知科学研究所; 德克萨斯大学奥斯汀分校语言学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出生成对齐诊断阶梯,区分语音语言模型的决策规则失配与读出覆盖限制,发现状态解码比生成准确率高27.8点,无标签logit校正可提升生成准确率,明确性能损失来源。

AI 中文摘要

语音语言模型越来越多地基于提示答案的准确性在副语言任务上进行评估,但答案准确性综合了从音频到答案计算不同阶段的失败。我们提出了一种与生成对齐的诊断阶梯,该阶梯比较生成的答案、选项 logits(对数几率)、这些 logits 的仿射读出,以及同一答案标记处隐藏状态的线性读出。连续的差异可区分端点、决策规则和读出覆盖差距。在五个系统和两个情感语料库中,状态解码平均比生成高出 27.8 个准确率点,且在所有 10 种条件下,决策规则和读出覆盖差距均为正。无标签 logit 校正可在每种条件下提高生成准确率,表明部分决策规则差距是可操作的。在秩匹配比较中,原生读出之外的情感信息可泛化到未见过的说话者,且在对测量的声学描述符进行控制后仍保留,但替换选定的读出外部方向通常对生成的答案影响很小。这些结果将信息可用性与行为使用区分开来,并将性能损失定位在决策规则和状态到答案的读出环节。

英文摘要

Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token. Successive differences separate endpoint, decision-rule, and readout-coverage gaps. Across five systems and two emotion corpora, state decoding exceeds generation by 27.8 accuracy points on average, and both the decision-rule and readout-coverage gaps are positive in all ten conditions. A label-free logit correction improves generated accuracy in every condition, showing that part of the decision-rule gap is actionable. In rank-matched comparisons, emotion information outside the native readout generalizes to held-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout-external directions usually has little effect on emitted answers. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state-to-answer readout.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑