arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

农业领域模型无关且语言无关的语音管道改进

Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain

Aakash Singh, Lakshmi Pedapudi, Chandrashekar M S, Sanyam Singh, Naga Ganesh, Vineet Singh

arXiv 2609.20504首次发表:更新:

发表机构

Digital Green(数字绿色组织)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种模型无关且语言无关的模块化语音管道,通过门控音频增强、说话人分离与选择、领域感知校正及质量门控,在不微调ASR模型的情况下,将农业语音转录的WER相对降低16-23%(云模型)和5%(设备端模型),多说话人场景降幅更大。

AI 中文摘要

FarmerChat是Digital Green为小农户提供的AI驱动的农业咨询助手,农户可以通过文本、语音或照片以他们自己的语言访问它。语音对这一人群来说是一个关键渠道,然而现场录制的语音对通用自动语音识别(ASR)来说具有挑战性,因为录音中经常包含机器噪音、背景媒体、竞争说话者和特定领域的农业词汇。这些条件对承载农户查询意义的作物、害虫、化学品和数量术语影响尤为严重。我们提出了一种模块化、模型无关的管道,用于在不微调或替换底层ASR模型的情况下改进FarmerChat中的ASR质量。该管道结合了门控音频增强、说话人分离和目标说话人选择、ASR、使用加权农业词典的领域感知校正,以及用于检测不可靠转录的质量门控。只有说话人分离阶段进行了微调;所有其他阶段使用常见接口背后的现成模型。我们在印地语、泰卢固语和奥里亚语的人工标注FarmerChat录音上评估了该管道,使用词错误率(WER)和领域加权错误率,后者对农业术语赋予更大重要性。最大的改进出现在多说话人录音上,其中目标说话人选择防止竞争语音进入转录。在整个语料库中,该管道在三个云ASR模型上将WER相对降低了16-23%,在设备端模型上降低了5%。在多说话人录音上,云模型的降低幅度为32-42%,设备端模型为16%。所有报告的降低在统计上均显著。这些结果表明,有针对性的预处理、说话人选择和领域感知的后处理可以在保留底层ASR模型的同时显著改善农业语音转录。

英文摘要

FarmerChat is Digital Green's AI-powered agricultural advisory assistant for smallholder farmers, who access it in their own language through text, voice, or photographs. Voice is a critical channel for this population, yet field-recorded speech is challenging for general-purpose automatic speech recognition (ASR) because recordings frequently contain machinery noise, background media, competing speakers, and domain-specific agricultural vocabulary. These conditions disproportionately affect crop, pest, chemical, and quantity terms that carry the meaning of a farmer's query. We present a modular, model-agnostic pipeline for improving ASR quality in FarmerChat without fine-tuning or replacing the underlying ASR model. The pipeline combines gated audio enhancement, speaker diarization and target-speaker selection, ASR, domain-aware correction using a weighted agricultural lexicon, and a quality gate for detecting unreliable transcripts. Only the diarization stage is fine-tuned; all other stages use off-the-shelf models behind common interfaces. We evaluate the pipeline on human-annotated FarmerChat recordings in Hindi, Telugu, and Odia using word error rate (WER) and a domain-weighted error rate that gives greater importance to agricultural terminology. The largest improvements occur on multi-speaker recordings, where target-speaker selection prevents competing speech from entering the transcript. Across the full corpus, the pipeline reduces WER by 16-23% relative on three cloud ASR models and by 5% on an on-device model. On multi-speaker recordings, the reductions are 32-42% for the cloud models and 16% for the on-device model. All reported reductions are statistically significant. These results show that targeted preprocessing, speaker selection, and domain-aware post-processing can substantially improve agricultural speech transcription while preserving the underlying ASR model.

Comments20 tables, 11 figures, 23 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑