Key points
- Apache-licensed open video model family spanning text-to-video, image-to-video and video editing tasks.
- The 1.3 billion parameter variant runs in approximately 8 gigabytes of VRAM, within reach of student hardware.
- Topped the VBench leaderboard at release, ahead of closed-source competitors.
Summary
The Wan model family, presented in a paper published in March 2025, is an Apache-licensed open video generation system covering text-to-video, image-to-video and editing tasks. The smallest variant (1.3 billion parameters) runs within approximately 8 gigabytes of VRAM, a threshold accessible on student-grade hardware. At release the Wan models led the VBench benchmark, outperforming closed-source competitors. Selected for the Knowledge Base over HunyuanVideo and Mochi on the basis of its distinctive educator value: the open licence and hardware accessibility make it the open video model a teaching lab can actually deploy.
Related items
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- AnimateDiff: Animate Your Personalized Text-to-Image Models without Specific Tuning
Source
Source: arXiv ↗ (Research)
Cite this item
arXiv (2025). ‘Wan: Open and Advanced Large-Scale Video Generative Models’, arXiv. Available at: https://arxiv.org/abs/2503.20314
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.