[论文] An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Mod…

## 论文概要 **研究领域**: CV **作者**: Dengyang Jiang, Ruoyi Du, ...

论文概要

研究领域: CV 作者: Dengyang Jiang, Ruoyi Du, Zhennan Chen et al. (13 authors) 发布时间: 2026-08-17 arXiv: 2608.16887

中文摘要

本文研究生成建模中一个日益重要的主题:像素空间扩散模型。尽管已有众多研究,但大多聚焦于小规模或类别条件设置,训练出能与成熟潜空间模型匹敌的像素空间模型的实用方案仍然难以捉摸。通过全面的实证研究,我们首先观察到直接在像素空间进行大规模预训练的收敛速度明显慢于潜空间。这促使我们提出潜空间到像素空间的策略:先在潜空间高效获取生成先验,然后在后训练阶段过渡到像素空间。我们系统研究了控制这一过渡的关键设计选择,包括权重初始化、数据组成、预测目标、解码器架构和噪声调度,并确定了一个实用方案,使得到的像素空间模型能够匹敌或超越其潜空间对应物,同时实现3.18到4.75倍的端到端推理加速。

原文摘要

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weigh…

自动采集于 2026-08-19

#论文 #arXiv #CV #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1