Skip to main content

Research · Technical

High-Resolution Image Synthesis with Latent Diffusion Models

CVPR 2022 (IEEE/CVF) · Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B. · Jun 2022

Key points

  1. Latent diffusion models run the denoising process in a compressed latent space, dramatically reducing the compute cost of high-resolution image synthesis.
  2. The architecture became the foundation of Stable Diffusion, the open-source model that is now the most widely deployed image-generation system in creative industries.
  3. The LDM architecture underlies Stable Diffusion and its derivatives, making it the technical basis of the image-generation tools most widely deployed in animation pre-production.

Summary

This paper introduces latent diffusion models (LDMs), which perform the iterative denoising process in a learned latent space rather than pixel space, making high-resolution image synthesis computationally tractable. The architecture became the technical foundation of Stable Diffusion and its derivatives, the dominant open-source image-generation systems now embedded in animation pre-production workflows including concept art, storyboarding, and moodboarding. For animation educators, this is the primary text behind the class of generative tools most likely to appear in student and studio practice.

Source

Source: CVPR 2022 (IEEE/CVF) ↗ (Research)

Cite this item

Rombach, R., Blattmann, A., Lorenz, D., Esser, P. & Ommer, B. (2022). ‘High-Resolution Image Synthesis with Latent Diffusion Models’, CVPR 2022 (IEEE/CVF). Available at: https://arxiv.org/abs/2112.10752

Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.