论文概要
研究领域: ML 作者: Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer, Roger Wattenhofer 发布时间: 2026-08-19 arXiv: 2608.19141
中文摘要
基于残差向量量化(RVQ)的神经音频编解码器已成为基于token的通用音频生成的主要离散表示,然而从粗编解码token重新合成高质量音频仍然是一个开放问题,并限制了每个生成它们的系统的保真度。先前工作将重新合成框架为离散token预测和连续回归之间的选择。我们认为这种二分法不完整,并引入几何迭代检索,一种使用RVQ层层次结构本身作为连续码本空间中自然迭代分解的范式。我们的方法不是在离散词汇上分类或回归到单一目标向量,而是在码本的几何空间中进行对比检索。我们在语音和音乐的编解码器恢复任务上评估我们的方法,并显示对单次token预测和一步回归基线的改进。
原文摘要
Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every system that generates them. Prior work has framed resynthesis as a choice between discrete token prediction and continuous regression. We argue that this dichotomy is incomplete and introduce geometric iterative retrieval, a paradigm that uses the RVQ layer hierarchy itself as a natural iterative decomposition in continuous codebook space. Rather than classifying over discrete vocabularies or regressing to a single target vector, our method performs contrastive retrieval in the codebook’s geometric space. We evaluate ou…
— 自动采集于 2026-08-21
#论文 #arXiv #ML #小凯
