Florence: A New Foundation Model for Computer Vision
Lu Yuan Affiliation: Microsoft Cloud and AI Correspondence to: Dongdong Chen Affiliation: Microsoft Cloud and AI Yi-Ling Chen Affiliation: Microsoft Cloud and AI Noel Codella Affiliation: Microsoft Cloud and AI Xiyang Dai Affiliation: Microsoft Cloud and AI Jianfeng Gao Affiliation: Microsoft Research Redmond Houdong Hu Affiliation: Microsoft Cloud and AI Xuedong Huang Affiliation: Microsoft Cloud and AI Boxin Li Affiliation: Microsoft Cloud and AI Chunyuan Li Affiliation: Microsoft Research Redmond Ce Liu Affiliation: Microsoft Cloud and AI Mengchen Liu Affiliation: Microsoft Cloud and AI Zicheng Liu Affiliation: Microsoft Cloud and AI Yumao Lu Affiliation: Microsoft Cloud and AI Yu Shi Affiliation: Microsoft Cloud and AI Lijuan Wang Affiliation: Microsoft Cloud and AI Jianfeng Wang Affiliation: Microsoft Cloud and AI Bin Xiao Affiliation: Microsoft Cloud and AI Zhen Xiao Affiliation: Microsoft Cloud and AI Jianwei Yang Affiliation: Microsoft Research Redmond Michael Zeng Affiliation: Microsoft Cloud and AI Luowei Zhou Affiliation: Microsoft Cloud and AI Pengchuan Zhang Affiliation: Microsoft Research Redmond
Abstract
Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, which are trained on diverse, large-scale dataset and can be adapted to a wide range of downstream tasks, are critical for this mission to solve real-world computer vision applications. While existing vision foundation models such as CLIP (Radford et al. 2021), ALIGN (Jia et al. 2021), and Wu Dao 2.0 (Wud) focus mainly on mapping images and textual representations to a cross-modal shared representation, we introduce a new computer vision foundation model, Florence, to expand the representations from coarse (scene) to fine (object), from static (images) to dynamic (videos), and from RGB to multiple modalities (caption, depth). By incorporating universal visual-language representations from Web-scale image-text data, our Florence model can be easily adapted for various computer vision tasks, such as classification, retrieval, object detection, VQA, image caption, video retrieval and action recognition. Moreover, Florence demonstrates outstanding performance in many type
原文 arXiv:2111.11432;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2111.11432v1