韵律训练表示是否有助于超越可训练融合?一项与冻结HuBERT的参数匹配研究
Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT
浏览论文内容
中文总结 AI 辅助
本研究通过参数匹配的冻结HuBERT实验,发现韵律辅助表示虽被模型依赖,但相比可训练融合并未带来显著的词错误率改进。
中文摘要 AI 辅助
显式的韵律线索可能有助于自发语音的自动语音识别(ASR),但辅助表示通常需要额外的可训练组件,这使得增益是来自辅助信息还是融合机制变得不清楚。我们使用冻结的HuBERT主干网络和一个64维表示来解决这个问题,该表示经过训练以预测对数基频(log F0)、浊音度、Delta log F0、对数能量和频谱倾斜。我们比较了冻结主干网络识别器(基线)、具有零辅助输入的可训练融合(空)以及提供学习表示的相同融合(学习)。在Buckeye、Switchboard和AMI IHM数据集上,与基线相比,空将词错误率(WER)降低了0.71-1.45个百分点,而学习与空的差异为+0.07、-0.09和+0.00个百分点,差异不显著。然而,在推理时移除或不匹配该表示会增加学习的WER。因此,学习依赖于该表示,但与参数匹配的对照相比,没有显示出可测量的增量WER收益。
英文摘要
Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.
发表机构
- University of Arizona(亚利桑那大学)
- Speak(Speak公司)
机构由 AI 辅助整理,请以论文原文为准。