TY - JOUR TI - LAION-5B (open image-text training dataset) AU - Schuhmann, C. AU - Beaumont, R. AU - Vencu, R. AU - Gordon, C. AU - Wightman, R. AU - Cherti, M. AU - Coombes, T. AU - Katta, A. AU - Mullis, C. AU - Wortsman, M. AU - et al. PY - 2022 DA - 2022/10// JO - NeurIPS 2022 Datasets and Benchmarks AB - Schuhmann et al. (2022) released LAION-5B, an open dataset of approximately 5.85 billion image-text pairs assembled by filtering Common Crawl web data using CLIP similarity scores. Stable Diffusion and the majority of open generative image models were trained on LAION-5B subsets, making it the concrete training-data substrate beneath the open generative-AI ecosystem. The dataset is named in the Andersen v Stability AI copyright litigation as the dataset at issue, creating a direct link between the technical and legal entries in the repository. Teaching data provenance, consent, scraping ethics and copyright exposure in AI image generation requires understanding LAION-5B specifically. KW - training-data KW - ip-and-copyright KW - generative-ai UR - https://arxiv.org/abs/2210.08402 LA - en ER -