Key points
- DreamBooth fine-tunes an entire diffusion model on three to five subject images, binding a unique token to that subject for generation in novel contexts.
- It is the first practical character-lock-in fine-tuning method and has been widely productised for consistent character generation in animation pre-production workflows.
- DreamBooth and LoRA represent the two main fine-tuning paradigms for character consistency in generative animation: full subject fine-tuning versus low-rank adaptation.
Summary
Ruiz et al. (2022, published CVPR 2023) introduce DreamBooth, a fine-tuning method that binds a unique text identifier to a specific subject by fine-tuning all weights of a diffusion model on a small reference image set, with a prior-preservation loss that prevents language drift. The result allows the subject to be synthesised in arbitrary contexts, poses, and scenes while maintaining identity. DreamBooth was the first practically accessible character-lock-in method and has been widely productised in animation pre-production pipelines for consistent character generation.
Related items
- LoRA: Low-Rank Adaptation of Large Language Models
- High-Resolution Image Synthesis with Latent Diffusion Models
Source
Source: CVPR 2023 (IEEE/CVF) ↗ (Research)
Cite this item
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M. & Aberman, K. (2023). ‘DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation’, CVPR 2023 (IEEE/CVF). Available at: https://arxiv.org/abs/2208.12242
Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.