Key points
- Introduces the multimodal diffusion transformer and rectified flow training behind Stable Diffusion 3.
- Authored by the team that founded Black Forest Labs and built the FLUX model family.
- Demonstrates predictable scaling laws for image synthesis quality from 800 million to 8 billion parameters.
Summary
Published in March 2024, this paper introduces the multimodal diffusion transformer (MMDiT) architecture and rectified flow training methodology that underpin Stable Diffusion 3. The authors subsequently founded Black Forest Labs and applied the same approach to the FLUX model family, making this the conceptual lineage anchor for the image generation architectures currently taught across image-generation curricula. The paper demonstrates predictable scaling behaviour across a large parameter range, a finding with direct implications for understanding the capabilities and limits of the models students use.
Related items
- Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)
- Wan: Open and Advanced Large-Scale Video Generative Models
- High-Resolution Image Synthesis with Latent Diffusion Models
Source
Source: arXiv ↗ (Research)
Cite this item
Esser, P. & Stability AI research team (2024). ‘Scaling Rectified Flow Transformers for High-Resolution Image Synthesis’, arXiv. Available at: https://arxiv.org/abs/2403.03206
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.