Key points
- Adds spatial conditioning such as pose, depth and edges to pretrained text to image diffusion models.
- Enables pose controlled character generation, now core to AI character design teaching.
- Among the most cited creative AI papers, with over 5,000 citations.
Summary
ControlNet introduces a neural network architecture that adds spatial conditioning inputs (pose skeletons, depth maps, edge maps, and other structural signals) to pretrained text-to-image diffusion models without requiring full retraining. The result is controllable image generation where the spatial layout and character pose can be specified precisely, a capability that has become foundational in AI character design workflows. With over 5,000 citations, ControlNet is a landmark of the generative image field and appears in virtually every node-based workflow taught in animation and concept art courses today.
Related items
- AnimateDiff: Animate Your Personalized Text-to-Image Models without Specific Tuning
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- High-Resolution Image Synthesis with Latent Diffusion Models
Source
Source: IEEE/CVF International Conference on Computer Vision (ICCV 2023) ↗ (Research)
Cite this item
Zhang, L., Rao, A. & Agrawala, M. (2023). ‘Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)’, IEEE/CVF International Conference on Computer Vision (ICCV 2023). Available at: https://openaccess.thecvf.com/content/ICCV2023/html/Zhang_Adding_Conditional_Control_to_Text-to-Image_Diffusion_Models_ICCV_2023_paper.html
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.