Skip to main content

Research · Technical

Learning Transferable Visual Models From Natural Language Supervision (CLIP)

ICML 2021 · Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I. · Feb 2021

Key points

  1. CLIP learns a joint image-text embedding from large-scale web pairs, enabling language to guide image models.
  2. It is the semantic bridge that makes text prompts steer Stable Diffusion, DALL-E guidance and much ControlNet conditioning.
  3. Prompt-to-image teaching is not explicable without it.

Summary

Radford et al. (2021) trained a model to align image and text representations in a shared embedding space using contrastive learning on 400 million web-scraped image-caption pairs. The resulting CLIP model can match images to natural-language descriptions with strong zero-shot performance. CLIP is the component that allows text prompts to steer image-generation models: Stable Diffusion, DALL-E and CLIP-guided diffusion all depend on its joint embedding to translate a prompt into the visual direction for the generative process. Explaining prompt-to-image generation to animation students requires understanding CLIP's role as the semantic bridge between language and image space.

Source

Source: ICML 2021 ↗ (Research)

Cite this item

Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G. & Sutskever, I. (2021). ‘Learning Transferable Visual Models From Natural Language Supervision (CLIP)’, ICML 2021. Available at: https://arxiv.org/abs/2103.00020

Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.