Key points
- Video LDM extends latent diffusion to video by inserting temporal attention layers into a pretrained image diffusion model, enabling high-resolution video synthesis without training from scratch.
- The temporal-attention architecture introduced here is the direct technical ancestor of the latent video generation category, including AnimateDiff and subsequent production video-generation tools.
- Video LDM's temporal-attention insertion design became the template for the latent video generation category, including AnimateDiff, now entering visual-effects and motion-graphics workflows.
Summary
Blattmann et al. (2023) extend the latent diffusion model architecture to video by inserting temporal attention and 3D convolution layers into a pretrained image LDM, aligning the temporal dimension to produce temporally coherent high-resolution video. Published at CVPR 2023, the paper is the architectural ancestor of the latent video generation category. Admitted as a distinct foundational entry per owner decision QL3, it provides the conceptual origin for video-generation tools, including AnimateDiff, that are entering animation and motion-graphics production workflows.
Related items
- AnimateDiff: Animate Your Personalized Text-to-Image Models without Specific Tuning
- High-Resolution Image Synthesis with Latent Diffusion Models
Source
Source: CVPR 2023 (IEEE/CVF) ↗ (Research)
Cite this item
Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S. & Kreis, K. (2023). ‘Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models’, CVPR 2023 (IEEE/CVF). Available at: https://arxiv.org/abs/2304.08818
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.