论文概要
研究领域: CV 作者: Neta Shaul, Chao Liu, Arash Vahdat, Julius Berner 发布时间: 2026-07-28 arXiv: 2607.26004
中文摘要
视频扩散或流模型中的生成计算成本高昂,原因在于缓慢且迭代的采样过程。当前最先进的加速方法严重依赖变分分数蒸馏(VSD)和对抗损失将扩散模型蒸馏为少步生成器。尽管实现了高质量视频生成,这些训练损失 notoriously难以优化且遭受模式崩溃,导致视频多样性丧失和运动不足。本文中,我们引入了并行解码蒸馏(PDD),一种简化的、可扩展的基于轨迹的蒸馏方法,用于扩散和流匹配模型的快速推理。我们的架构和训练程序与任何预训练模型兼容,并支持使用 varying数量的函数评估(NFE)进行采样。PDD通过每次网络评估预测多个去噪步骤来加速生成。从概念上讲,它学习平均速度表征,而不使用JVPs或有限差分近似来回归其导数。我们的方法在LTX-2.3文本到视频/音频、Wan 14B文本到视频和Qwen-Image文本到图像上以4-8 NFE实现了最先进的性能。此外,PDD在生成视频多样性方面表现出显著改进。
原文摘要
Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion. In this paper, we introduce Parallel Decoding Distillation (PDD), a simplified and scalable trajectory-based distillation method for fast inference of diffusion and flow matching models. Our architecture and training procedure are compatible with any pre-trained model and support sampling with a varying n…
— 自动采集于 2026-07-30
#论文 #arXiv #CV #小凯
