LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Christoph Schuhmann LAION Richard Vencu LAION Gentec Data Romain Beaumont LAION Robert Kaczmarczyk LAION Technical University of Munich Clayton Mullis LAION Aarush Katta LAION Theo Coombes LAION Jenia Jitsev LAION Juelich Supercomputing Center (JSC) Research Center Juelich (FZJ) Aran Komatsuzaki LAION Georgia Institute of Technology EleutherAI
Abstract
Multi-modal language-vision models trained on hundreds of millions of image-text pairs (e.g. CLIP, DALL-E) gained a recent surge, showing remarkable capability to perform zero- or few-shot learning and transfer even in absence of per-sample labels on target image data. Despite this trend, to date there has been no publicly available datasets of sufficient scale for training such models from scratch. To address this issue, in a community effort we build and release for public LAION-400M, a dataset with CLIP-filtered 400 million image-text pairs, their CLIP embeddings and kNN indices that allow efficient similarity search.111Project page: https://laion.ai/laion-400-open-dataset/
中文速览
训练像CLIP、DALL-E这样强大的多模态视觉语言模型需要数亿规模的图文对数据,但此前从未有过公开可用的同等量级数据集。为此,研究者们以社区协作方式从Common Crawl网络爬取海量图文对,并利用CLIP模型计算图文相似度进行质量过滤,最终构建并公开发布了包含4亿图文对的LAION-400M数据集,同时附带CLIP嵌入向量和近邻检索索引。实验表明,仅用其中约720万张图片训练一个DALL-E架构模型就已能生成质量可观的图像,验证了数据集的有效性。这一数据集的开放发布意义重大——它填补了学术界与工业界在大规模图文预训练数据上的鸿沟,让更广泛的研究者得以在不依赖私有数据的情况下,从头复现和探索顶尖视觉语言模型。
原文 arXiv:2111.02114;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2111.02114v1