Skip to main content

Research · Technical

Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models

CVPR 2023 (IEEE/CVF) · Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., Kreis, K. · Apr 2023

Key points

  1. Video LDM extends latent diffusion to video by inserting temporal attention layers into a pretrained image diffusion model, enabling high-resolution video synthesis without training from scratch.
  2. The temporal-attention architecture introduced here is the direct technical ancestor of the latent video generation category, including AnimateDiff and subsequent production video-generation tools.
  3. Video LDM's temporal-attention insertion design became the template for the latent video generation category, including AnimateDiff, now entering visual-effects and motion-graphics workflows.

Summary

Blattmann et al. (2023) extend the latent diffusion model architecture to video by inserting temporal attention and 3D convolution layers into a pretrained image LDM, aligning the temporal dimension to produce temporally coherent high-resolution video. Published at CVPR 2023, the paper is the architectural ancestor of the latent video generation category. Admitted as a distinct foundational entry per owner decision QL3, it provides the conceptual origin for video-generation tools, including AnimateDiff, that are entering animation and motion-graphics production workflows.

Source

Source: CVPR 2023 (IEEE/CVF) ↗ (Research)

Cite this item

Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S. & Kreis, K. (2023). ‘Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models’, CVPR 2023 (IEEE/CVF). Available at: https://arxiv.org/abs/2304.08818

Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.