论文概要
研究领域: ML 作者: Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos 发布时间: 2026-08-19 arXiv: 2608.19174
中文摘要
本技术报告描述了我们赢得AES AIMLA 2025挑战赛的提交方案,该挑战赛要求通过声音模仿来查询音效。我们研究了两种互补的微调策略:使用冻结的预训练CED编码器进行对比学习,以及使用MobileNetV3编码器进行带有半难负样本的联合对比-三元组学习。本报告已更新以包含挑战赛后发布的详细信息。
原文摘要
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.
— 自动采集于 2026-08-21
#论文 #arXiv #ML #小凯
