arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01737cs.CL

SpeakPay:针对低资源尼泊尔语金融语音识别的Whisper领域自适应LoRA微调

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

Biraj Subedi

首次发表
浏览论文内容

中文总结 AI 辅助

针对视障用户无法使用尼泊尔图形化移动支付的问题,提出SpeakPay语音数字钱包,通过LoRA微调Whisper实现低资源尼泊尔语金融语音识别,大幅降低词错误率、提升交易成功率。

中文摘要 AI 辅助

尼泊尔的移动支付应用以图形界面为主,视障用户大多无法使用。本文提出了一款以语音为核心的数字钱包SpeakPay,并阐述了核心技术贡献:针对低资源金融语音识别的领域适配控制研究。我们推出了NepFinSpeech-403,这是一个包含403句尼泊尔语金融语音指令的数据集(涵盖发送、充值、余额查询操作,涉及237个唯一数字),并使用LoRA对Whisper large-v2进行微调。在保留的测试集上,经领域适配的模型将词错误率从129.95%(零样本基线)降至42.58%,相对降低67.2%,并将天城文数字识别准确率从0.0%提升至73.9%。我们发现词级指标低估了实际任务级影响:领域适配将交易成功率从1.67%提升至33.33%,增幅约20倍。该改进在单句层面(符号检验,p<10^-17)和所有指令类型中均保持一致。数据效率分析显示,仅需100条领域特定语句即可将零样本词错误率减半,性能在约300个样本时趋于平稳。错误分析揭示了系统性数字混淆模式(零插入/删除、前缀幻觉),这些模式是剩余交易失败的主要原因。训练后的系统已作为公开可访问的语音优先网络应用部署,所有代码、数据集、模型权重及本文均已发布在该httpsURL。

英文摘要

Mobile payment applications in Nepal are graphically mediated and largely inaccessible to visually impaired users. This paper presents SpeakPay, a voice-first digital wallet, and documents the central technical contribution: a controlled study of domain adaptation for low-resource financial speech recognition. We introduce NepFinSpeech-403, a 403-utterance dataset of Nepali financial voice commands (send, load, and balance operations spanning 237 unique numerals), and fine-tune Whisper large-v2 with LoRA. On the held-out test set, the domain-adapted model reduces Word Error Rate from 129.95% (zero-shot baseline) to 42.58% --- a 67.2% relative reduction --- and improves Devanagari numeral recognition accuracy from 0.0% to 73.9%. We find that word-level metrics understate the practical task-level impact: domain adaptation improves the Transaction Success Rate from 1.67% to 33.33%, a roughly 20x gain. The improvement is consistent at the individual-utterance level (sign test, $p < 10^{-17}$) and across all command types. A data efficiency analysis shows that as few as 100 domain-specific utterances are sufficient to halve the zero-shot WER, with performance plateauing around 300 examples. Error analysis reveals systematic numeral confusion patterns (zero insertion/deletion, prefix hallucination) that account for the majority of remaining transaction failures. The trained system is deployed as a publicly accessible voice-first web application. All code, dataset, model weights, and this paper are released at https://github.com/subedibiraj/speakpay.

补充信息

↑