Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data
利用音频大语言模型过滤语音到语音训练数据
机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院) ; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)
AI总结 提出Rank-to-Distill策略,训练音频大语言模型直接从语音对判断保留/丢弃,过滤噪声数据,提升端到端语音翻译性能。
Comments Accepted to INTERSPEECH 2026