Key points
- Sora scales a diffusion transformer to video by treating clips as sequences of spacetime patches, generating up to a minute of high-definition video.
- Its outputs show emergent 3D consistency and object permanence without explicit geometric supervision, suggesting learned world-model behaviour.
- Its February 2024 release reset the discourse on text-to-video in animation and visual-effects practice and education.
Summary
OpenAI's Sora technical report, published in February 2024, describes a diffusion transformer model that generates high-definition video up to one minute long by treating video clips as sequences of spacetime patches. The report documents emergent properties in Sora's outputs, including 3D consistency, object permanence, and plausible physical interactions, that were not the subject of explicit supervision. The February 2024 release had an immediate and measurable effect on discourse in the animation and visual-effects fields, resetting expectations about the near-term capability ceiling for text-to-video generation.
Related items
- Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
- Wan: Open and Advanced Large-Scale Video Generative Models
Source
Source: OpenAI (technical report) ↗ (Other)
Cite this item
OpenAI (technical report) (2024). ‘Video Generation Models as World Simulators (Sora technical report)’, OpenAI (technical report). Available at: https://openai.com/index/video-generation-models-as-world-simulators/
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.