LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann1 §§°° Romain Beaumont1 §§°° Richard Vencu1,3,8 §§°° Cade Gordon2 §§°° Ross Wightman1§§ Mehdi Cherti 1,10§§ Theo Coombes1 Aarush Katta1 Clayton Mullis1 Mitchell Wortsman6 Patrick Schramowski1,4,5 Srivatsa Kundurthy1 Katherine Crowson1,8,9 Ludwig Schmidt6 °° Robert Kaczmarczyk1,7 °° Jenia Jitsev1,10 °° LAION1 UC Berkeley2 Gentec Data3 TU Darmstadt4 Hessian.AI5 University of Washington, Seattle6 Technical University of Munich7 Stability AI8 EleutherAI9 Juelich Supercomputing Center (JSC), Research Center Juelich (FZJ)10 §§ Equal first contributions, °° Equal senior contributions
Abstract
Groundbreaking language-vision architectures like CLIP and DALL-E proved the utility of training on large amounts of noisy image-text data, without relying on expensive accurate labels used in standard vision unimodal supervised learning. The resulting models showed capabilities of strong text-guided image generation and transfer to downstream tasks, while performing remarkably at zero-shot classification with noteworthy out-of-distribution robustness. Since then, large-scale language-vision models like ALIGN, BASIC, GLIDE, Flamingo and Imagen made further improvements. Studying the training and capabilities of such models requires datasets containing billions of image-text pairs. Until now, no datasets of this size have been made openly available for the broader research community. To address this problem and democratize research on large-scale multi-modal models, we present LAION-5B - a dataset consisting of 5.85 billion CLIP-filtered image-text pairs, of which 2.32B contain English language. We show successful replication and fine-tuning of foundational models like CLIP, GLIDE and Stable Diffusion using the dataset, and discuss further experiments enabled with an openly availabl
中文速览
研究人员长期苦于没有公开可用的超大规模图文数据集——CLIP、DALL-E等突破性模型背后动辄数十亿对的训练数据从未对外公开,导致前沿多模态研究被少数工业实验室垄断。为此,作者团队从Common Crawl网页存档出发,用CLIP模型对图片与配套alt-text的语义相关性进行过滤筛选,最终构建并公开发布了LAION-5B——一个包含58.5亿图文对的数据集,其中23.2亿条为英文,另有22.6亿条覆盖多语言,是迄今最大的公开图文数据集。利用该数据集,作者成功复现了OpenAI CLIP(ViT-L/14)、GLIDE以及Stable Diffusion等基础模型,性能与原版相当,证明了数据集的可用性。LAION-5B的公开不仅让学术界首次有机会在同等规模的数据上开展审计、去偏和创新研究,也为多语言低资源场景和文生图等方向打开了新空间,是推动多模态AI研究民主化的重要一步。
原文 arXiv:2210.08402;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2210.08402v1