Key points
- CLIP learns a joint image-text embedding from large-scale web pairs, enabling language to guide image models.
- It is the semantic bridge that makes text prompts steer Stable Diffusion, DALL-E guidance and much ControlNet conditioning.
- Prompt-to-image teaching is not explicable without it.
Summary
Radford et al. (2021) trained a model to align image and text representations in a shared embedding space using contrastive learning on 400 million web-scraped image-caption pairs. The resulting CLIP model can match images to natural-language descriptions with strong zero-shot performance. CLIP is the component that allows text prompts to steer image-generation models: Stable Diffusion, DALL-E and CLIP-guided diffusion all depend on its joint embedding to translate a prompt into the visual direction for the generative process. Explaining prompt-to-image generation to animation students requires understanding CLIP's role as the semantic bridge between language and image space.
Related items
- High-Resolution Image Synthesis with Latent Diffusion Models
- Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)
Source
Source: ICML 2021 ↗ (Research)
Cite this item
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G. & Sutskever, I. (2021). ‘Learning Transferable Visual Models From Natural Language Supervision (CLIP)’, ICML 2021. Available at: https://arxiv.org/abs/2103.00020
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.