Skip to main content

Research · Technical

LAION-5B (open image-text training dataset)

NeurIPS 2022 Datasets and Benchmarks · Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. · Oct 2022

Key points

  1. LAION-5B is the open 5-billion image-text dataset used to train Stable Diffusion and most open models.
  2. It is the concrete link between how models learn and the training-data copyright cases the repository already holds.
  3. It is central to teaching data provenance, consent and scraping debates.

Summary

Schuhmann et al. (2022) released LAION-5B, an open dataset of approximately 5.85 billion image-text pairs assembled by filtering Common Crawl web data using CLIP similarity scores. Stable Diffusion and the majority of open generative image models were trained on LAION-5B subsets, making it the concrete training-data substrate beneath the open generative-AI ecosystem. The dataset is named in the Andersen v Stability AI copyright litigation as the dataset at issue, creating a direct link between the technical and legal entries in the repository. Teaching data provenance, consent, scraping ethics and copyright exposure in AI image generation requires understanding LAION-5B specifically.

Source

Source: NeurIPS 2022 Datasets and Benchmarks ↗ (Research)

Cite this item

Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M. & et al. (2022). ‘LAION-5B (open image-text training dataset)’, NeurIPS 2022 Datasets and Benchmarks. Available at: https://arxiv.org/abs/2210.08402

Your reference manager can also read this page directly: with the Zotero (or Mendeley) browser connector installed, save it straight to your library. Whole-collection exports: RIS, BibTeX, CSL-JSON.