arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NOPE-HYPE:一种面向多样声学环境的鲁棒语音转文本的结构化仿真工作流

NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic Environments

Niramay M. Patel, Bibek Behera, Raksha Sharma

arXiv 2609.10058首次发表:更新:

发表机构

IISER Bhopal; IIT Bombay; IIT Roorkee(印度科学教育与研究院博帕尔分院; 印度理工学院孟买分校; 印度理工学院鲁尔基分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

NOPE-HYPE提出结合可控环境模拟器、PSD覆盖缩减和超参数搜索的结构化训练工作流,提升语音转文本在多样声学环境中的鲁棒性。

AI 中文摘要

鲁棒的语音转文本翻译系统应在多样声学条件下可靠运行,然而实际流程缺乏可控工具进行系统化环境探索。大型语音模型对未见声学条件仍敏感,因为训练数据很少覆盖真实环境的全部范围。我们提出NOPE-HYPE,一种结构化训练工作流,结合可控环境模拟器、基于功率谱密度(PSD)模板的覆盖最优环境缩减,以及针对模拟器旋钮的小型可解释超参数搜索。我们表明,模拟器生成的噪声在Whisper和SeamlessM4T模型上实现了与平衡真实噪声训练相当的性能,提供了原则性的环境原型集,并从结构化的27次运行超参数扫描中识别出实用的默认模拟器配置。

英文摘要

Robust speech-to-text translation systems should perform reliably across diverse acoustic conditions, yet practical pipelines lack controllable tools for systematic environment exploration. Large speech models remain sensitive to unseen acoustic conditions, as training data rarely cover the full range of real environments.We present NOPEHYPE, a structured training workflow that combines a controllable environment simulator, coverage-optimal environment reduction on Power Spectral Density (PSD) templates, and a small, interpretable hyperparameter search over simulator knobs. We show that simulator-generated noise achieves performance comparable to balanced realnoise training across Whisper and SeamlessM4T models, provide principled environment prototype sets, and identify practical default simulator configurations from a structured 27-run hyperparameter sweep.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑