Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers

Submitted ⁨⁨1⁩ ⁨year⁩ ago⁩ by ⁨Even_Adder@lemmy.dbzer0.com⁩ to ⁨stable_diffusion@lemmy.dbzer0.com⁩

https://crowsonkb.github.io/hourglass-diffusion-transformers/

Abstract

We present the Hourglass Diffusion Transformer (HDiT), an image generative model that exhibits linear scaling with pixel count, supporting training at high-resolution (e.g. 1024×1024) directly in pixel-space. Building on the Transformer architecture, which is known to scale to billions of parameters, it bridges the gap between the efficiency of convolutional U-Nets and the scalability of Transformers. HDiT trains successfully without typical high-resolution training techniques such as multiscale architectures, latent autoencoders or self-conditioning. We demonstrate that HDiT performs competitively with existing models on ImageNet 2562, and sets a new state-of-the-art for diffusion models on FFHQ-10242.

Paper: arxiv.org/abs/2401.11605

Code: github.com/crowsonkb/k-diffusion

Project Page: …github.io/hourglass-diffusion-transformers/

source

Comments

Sort:hotnew top

tagginator@utter.online [bot] ⁨1⁩ ⁨year⁩ ago
New Lemmy Post: Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers (https://lemmy.dbzer0.com/post/12894500)
Tagging: #StableDiffusion
(Replying in the OP of this thread (NOT THIS BOT!) will appear as a comment in the lemmy discussion.)
I am a FOSS bot. Check my README: https://github.com/db0/lemmy-tagginator/blob/main/README.md
source