Key points
- Wav2Lip uses a pretrained lip-sync discriminator as a teacher signal to achieve accurate sync on in-the-wild video.
- It is speaker-agnostic and needs no per-subject fine-tuning, suiting dubbing and dialogue-replacement workflows.
- It is the open-source baseline named by all subsequent talking-face research.
Summary
Prajwal, Mukhopadhyay, Namboodiri and Jawahar (2020) introduced Wav2Lip, an audio-driven lip-synchronisation method that uses a pretrained lip-sync expert discriminator as a teacher to train a generator producing accurate mouth movements for any speaker given an audio track. The key insight is using a discriminator specifically trained to judge lip-sync quality rather than general visual realism, producing accurate synchronisation on in-the-wild video without per-subject fine-tuning. Wav2Lip is speaker-agnostic and directly applicable to dubbing and dialogue-replacement workflows in animation and localisation. It became the open-source baseline cited by all subsequent talking-face and audio-driven animation research.
Related items
Source
Source: ACM Multimedia 2020 ↗ (Research)
Cite this item
Prajwal, K. R., Mukhopadhyay, R., Namboodiri, V. P. & Jawahar, C. V. (2020). ‘Wav2Lip: A Lip Sync Expert Is All You Need’, ACM Multimedia 2020. Available at: https://arxiv.org/abs/2008.10010
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.