arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你何时可以剪枝你的网络?多语言语音解析中中间神经元的研究

When Can You Prune Your Network? A Study of Intermediate Neurons in Multilingual Speech Parsing

Minnie Kabra, Benjamin Lecouteux, Maximin Coavoux

arXiv 2610.11520首次发表:更新:

发表机构

Univ. Grenoble Alpes; CNRS; Grenoble INP; LIG(格勒诺布尔大学; 法国国家科学研究中心; 格勒诺布尔理工学院; 信息与信号处理实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多语言语音解析,提出移除中间神经网络单元的简化架构,参数减少12%且性能相当或更优,揭示中间神经元在预训练编码器冻结时可缩小表征差距,并评估了不同语言及相关因素的影响。

AI 中文摘要

端到端语音解析是近期提出的一项任务,旨在预测口语话语的转录文本和句法树。现有语音解析架构常使用中间神经网络。本研究考察中间神经网络(NN)在解析中的有效性,特别是其作用。我们提出一种更简单的端到端语音解析架构,移除这些中间NN单元,参数减少12%,在自动语音识别(ASR)和解析任务上的性能与现有方法相当或更优。我们证明,当预训练编码器被冻结时,中间NN单元有助于缩小表征差距。我们对法语、中低资源语言斯洛文尼亚语和奈及利亚英语(Naija)的语音解析进行全面评估,还研究了训练数据量和预训练语音编码器的中间层对语音解析的影响。

英文摘要

End-to-end speech parsing, a task recently proposed, consists in predicting both the transcription and the syntactic tree for a spoken utterance. Existing architectures for speech parsing often utilise intermediate neural networks. In this work, we examine the effectiveness of intermediate neural networks (NN) for parsing, and, specifically, what role do they play. We introduce a simpler end-to-end architecture for speech parsing, where we remove these intermediate NN units, reducing the parameters by 12%, while achieving comparable or better performance than prior method on both automatic speech recognition (ASR) and parsing. We demonstrate that intermediate NN units help reduce the representational gap when the pre-trained encoder is frozen. We do a comprehensive evaluation of speech parsing on French, and medium-low resource languages Slovenian and Naija. We further investigate the impact of the training data size and intermediate layers of the pretrained speech encoder on speech parsing.

Commentsto appear in Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑