arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于ASR评估的生成式与编码器大型语言模型:一项比较研究

Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu, Mickael Rouvier, Jane Wottawa, Richard Dufour

arXiv 2608.25574首次发表:更新:

发表机构

Idiap Research Institute; Avignon Université; Le Mans Université; Nantes Université(伊迪亚普研究所; 阿维尼翁大学; 勒芒大学; 南特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比生成式与编码器大型语言模型在ASR评估中的应用,发现合理配置的BERTScore、SemDist与人类判断相关性强,生成式LLMs可提升评估可解释性。

AI 中文摘要

自动语音识别(ASR)通常采用词错误率(WER)进行评估,但该指标无法很好地反映语义相似度。尽管基于嵌入的指标与人类判断的相关性更高,但基于编码器和解码器的大型语言模型(LLMs)在ASR评估中的各自作用仍未得到充分探索。本文对这两类模型在ASR评估中的应用展开比较研究。我们在不同LLMs、层及池化策略下分析BERTScore和SemDist,结果显示这两个指标在配置合理时均可与人类判断实现强相关性。针对解码器模型,我们在两种设置下研究生成式LLMs:通过提示进行假设的成对选择,以及直接定性错误分类。我们的研究结果表明,基于编码器的指标仍具强竞争力,而生成式LLMs在假设比较中表现突出,并提升了ASR评估的可解释性。

英文摘要

Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models (LLMs) remain underexplored. This paper presents a comparative study of both families for ASR evaluation. We analyze BERTScore and SemDist across different LLMs, layers, and pooling strategies, showing that both metrics can achieve strong correlation with human judgments when properly configured. For decoder models, we investigate generative LLMs in two settings: pairwise hypothesis selection via prompting and direct qualitative error classification. Our results show that encoder-based metrics remain highly competitive, while generative LLMs perform strongly in hypothesis comparison and improve the interpretability of ASR evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑