[论文] Finetuning Strategies for Querying Sounds by Vocal Imitation

## 论文概要 **研究领域**: ML **作者**: Aditya Bhattacharjee, Chri...

论文概要

研究领域: ML 作者: Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos 发布时间: 2026-08-19 arXiv: 2608.19174

中文摘要

本技术报告描述了我们赢得AES AIMLA 2025挑战赛的提交方案,该挑战赛要求通过声音模仿来查询音效。我们研究了两种互补的微调策略:使用冻结的预训练CED编码器进行对比学习,以及使用MobileNetV3编码器进行带有半难负样本的联合对比-三元组学习。本报告已更新以包含挑战赛后发布的详细信息。

原文摘要

This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.

自动采集于 2026-08-21

#论文 #arXiv #ML #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1