arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02735cs.CLcs.SD

为针对每位患者的构音障碍自动语音识别(ASR)选择参数高效微调(PEFT)变体:基于两个ASR基础模型的单说话人案例研究

Choosing a PEFT Variant for Per-Patient Dysarthric ASR: A Single-Speaker Case Study on Two ASR Bases

Bernard Muller, László Tóth, LaVonne Roberts

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对一名重度构音障碍的卒中后匈牙利男性说话人,对比七种PEFT变体在两款ASR基础模型上的表现,得出LoRA更具优势、5分钟音频采集可实现近一半CER降低量等结论。

中文摘要 AI 辅助

针对每位患者的适配器是构音障碍自动语音识别(ASR)的首选生产架构,但参数高效微调(PEFT)变体尚未在说话人依赖的、针对每位患者的场景中进行比较。我们开展了一项单说话人案例研究,针对一名卒中后匈牙利男性说话人(S1,共409条 utterance;听觉感知临床评估显示重度构音障碍),在两个生产基础模型(经匈牙利语微调的Whisper-large-v1,以及多语言Qwen3-ASR-1.7B检查点)上对比了七种LoRA系列方法(LoRA、QLoRA、AdaLoRA、DoRA、LoHA、VeRA、VB-LoRA)。注意力投影适配器在两个基础模型上均大幅降低了字符错误率(CER)。在三个随机种子下,配对自举检验未检测到LoRA与DoRA存在显著差异(p>0.5;在Whisper上的CER分别为13.86%/13.90%,在Qwen3-ASR上为28.10%/28.33%),因此我们采用更简单、成本更低的LoRA。实际4位(NF4)QLoRA在每个种子和两个基础模型上的表现均更差(CER分别为14.56%/30.09%),且在该规模下无内存节省;LoHA、VeRA、VB-LoRA和AdaLoRA的表现未达到LoRA系列水平,不过LoHA仍使Whisper上的CER相对降低了18.6%。在同一基础模型上,全微调的准确率更高(CER为11.43%),但一款适配前馈模块的115 MB LoRA在约为每位患者存储量3.7%的情况下,其CER与全微调仅相差0.66个百分点。6点注册网格显示,约5分钟的患者音频采集可实现零样本到30分钟的CER降低量的45.6%,在10分钟和30分钟时还会有进一步提升(注意:仅针对一名说话人、一种语言、重度卒中后构音障碍)。训练脚本和方案将在发表后以研究用途许可开源发布。

英文摘要

Per-patient adapters are the preferred production architecture for dysarthric automatic speech recognition (ASR), yet parameter-efficient fine-tuning (PEFT) variants have not been compared in the speaker-dependent, per-patient regime. We present a single-speaker case study comparing seven LoRA-family methods (LoRA, QLoRA, AdaLoRA, DoRA, LoHA, VeRA, VB-LoRA) on two production bases (Whisper-large-v3 with Hungarian fine-tuning, and a multilingual Qwen3-ASR-1.7B checkpoint) for one post-stroke Hungarian male speaker (S1, 409 utterances; severe dysarthria on auditory-perceptual clinical assessment). Attention-projection adapters substantially improve CER on both bases. Across three seeds, a paired bootstrap detects no significant LoRA-DoRA difference (p>0.5; 13.86/13.90 % CER on Whisper, 28.10/28.33 % on Qwen3-ASR), so we adopt the simpler, cheaper LoRA. Real 4-bit (NF4) QLoRA is worse on every seed and both bases (14.56/30.09 % CER) with no memory saving at this scale, and LoHA, VeRA, VB-LoRA and AdaLoRA do not reach the LoRA family, though LoHA still gives an 18.6 % relative CER reduction on Whisper. On the same base, full fine-tuning is more accurate (11.43 % CER), but a 115 MB LoRA that also adapts the feed-forward blocks reaches within 0.66 pp of it at approximately 3.7 % of the per-patient storage. A 6-point enrollment grid shows about 5 min of patient audio captures 45.6 % of the zero-shot-to-30-min CER reduction, with further gains at 10 and 30 min (caveat: one speaker, one language, severe post-stroke dysarthria). Training scripts and recipes will be released, source-available under a research-use licence, on publication.

发表机构

  • The Scott-Morgan Foundation(斯科特-摩根基金会)
  • Institute of Informatics, University of Szeged(塞格德大学信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑