Key points
- Latent diffusion models run the denoising process in a compressed latent space, dramatically reducing the compute cost of high-resolution image synthesis.
- The architecture became the foundation of Stable Diffusion, the open-source model that is now the most widely deployed image-generation system in creative industries.
- The LDM architecture underlies Stable Diffusion and its derivatives, making it the technical basis of the image-generation tools most widely deployed in animation pre-production.
Summary
This paper introduces latent diffusion models (LDMs), which perform the iterative denoising process in a learned latent space rather than pixel space, making high-resolution image synthesis computationally tractable. The architecture became the technical foundation of Stable Diffusion and its derivatives, the dominant open-source image-generation systems now embedded in animation pre-production workflows including concept art, storyboarding, and moodboarding. For animation educators, this is the primary text behind the class of generative tools most likely to appear in student and studio practice.
Related items
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)
- AnimateDiff: Animate Your Personalized Text-to-Image Models without Specific Tuning
- Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
Source
Source: CVPR 2022 (IEEE/CVF) ↗ (Research)
Cite this item
Rombach, R., Blattmann, A., Lorenz, D., Esser, P. & Ommer, B. (2022). ‘High-Resolution Image Synthesis with Latent Diffusion Models’, CVPR 2022 (IEEE/CVF). Available at: https://arxiv.org/abs/2112.10752
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.