Skip to main content

Research · Technical

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

arXiv · Esser, P., Stability AI research team · Mar 2024

Key points

  1. Introduces the multimodal diffusion transformer and rectified flow training behind Stable Diffusion 3.
  2. Authored by the team that founded Black Forest Labs and built the FLUX model family.
  3. Demonstrates predictable scaling laws for image synthesis quality from 800 million to 8 billion parameters.

Summary

Published in March 2024, this paper introduces the multimodal diffusion transformer (MMDiT) architecture and rectified flow training methodology that underpin Stable Diffusion 3. The authors subsequently founded Black Forest Labs and applied the same approach to the FLUX model family, making this the conceptual lineage anchor for the image generation architectures currently taught across image-generation curricula. The paper demonstrates predictable scaling behaviour across a large parameter range, a finding with direct implications for understanding the capabilities and limits of the models students use.

Source

Source: arXiv ↗ (Research)

Cite this item

Esser, P. & Stability AI research team (2024). ‘Scaling Rectified Flow Transformers for High-Resolution Image Synthesis’, arXiv. Available at: https://arxiv.org/abs/2403.03206

Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.