发表机构
School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦玛丽女王大学电子工程与计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AES AIMLA 2025挑战赛的音效语音查询任务,研究人员提出两种微调策略,即基于CED编码器的对比学习和基于MobileNetV3编码器的联合对比-三元组学习,其方案为挑战赛获胜方案。
AI 中文摘要
本技术报告介绍了我们在AES AIMLA 2025挑战赛中提交的获胜方案,该挑战赛任务是通过语音模仿查询音效。我们研究了两种互补的微调策略:使用冻结的预训练CED编码器的对比学习,以及使用MobileNetV3编码器的半难负样本联合对比-三元组学习。为便于后续参考,本报告已更新,纳入了挑战赛结束后发布的细节。
英文摘要
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.