TY - JOUR TI - Learning Transferable Visual Models From Natural Language Supervision (CLIP) AU - Radford, A. AU - Kim, J. W. AU - Hallacy, C. AU - Ramesh, A. AU - Goh, G. AU - Agarwal, S. AU - Sastry, G. AU - Askell, A. AU - Mishkin, P. AU - Clark, J. AU - Krueger, G. AU - Sutskever, I. PY - 2021 DA - 2021/02// JO - ICML 2021 AB - Radford et al. (2021) trained a model to align image and text representations in a shared embedding space using contrastive learning on 400 million web-scraped image-caption pairs. The resulting CLIP model can match images to natural-language descriptions with strong zero-shot performance. CLIP is the component that allows text prompts to steer image-generation models: Stable Diffusion, DALL-E and CLIP-guided diffusion all depend on its joint embedding to translate a prompt into the visual direction for the generative process. Explaining prompt-to-image generation to animation students requires understanding CLIP's role as the semantic bridge between language and image space. KW - generative-ai KW - image-generation KW - ai-literacy UR - https://arxiv.org/abs/2103.00020 LA - en ER -