arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过语音模仿查询声音的微调策略

Finetuning Strategies for Querying Sounds by Vocal Imitation

Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos

arXiv 2608.19174首次发表:更新:

发表机构

School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦玛丽女王大学电子工程与计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AES AIMLA 2025挑战赛的音效语音查询任务,研究人员提出两种微调策略,即基于CED编码器的对比学习和基于MobileNetV3编码器的联合对比-三元组学习,其方案为挑战赛获胜方案。

AI 中文摘要

本技术报告介绍了我们在AES AIMLA 2025挑战赛中提交的获胜方案,该挑战赛任务是通过语音模仿查询音效。我们研究了两种互补的微调策略:使用冻结的预训练CED编码器的对比学习,以及使用MobileNetV3编码器的半难负样本联合对比-三元组学习。为便于后续参考,本报告已更新,纳入了挑战赛结束后发布的细节。

英文摘要

This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑