Generalized Decoding for Pixel, Image, and Language
Xueyan Zou∗§, Zi-Yi Dou∗♯, Jianwei Yang∗‡♠‡absent♠\ddagger\spadesuit, Zhe Gan†, Linjie Li†, Chunyuan Li‡, Xiyang Dai†, Harkirat Behl‡ Jianfeng Wang†, Lu Yuan†, Nanyun Peng♯, Lijuan Wang†, Yong Jae Lee¶§¶§\P\S, Jianfeng Gao¶‡\P\ddagger § University of Wisconsin-Madison ♯ UCLA ‡ Microsoft Research at Redmond † Microsoft Cloud、AI ∗Equal Technical Contribution ¶¶\P Equal Advisory Contribution ♠♠\spadesuit Project Lead
Abstract
We present X-Decoder, a generalized decoding model that can predict pixel-level segmentation and language tokens seamlessly. X-Decoder takes as input two types of queries: ( $i$ ) generic non-semantic queries and ( $ii$ ) semantic queries induced from text inputs, to decode different pixel-level and token-level outputs in the same semantic space. With such a novel design, X-Decoder is the first work that provides a unified way to support all types of image segmentation and a variety of vision-language (VL) tasks. Further, our design enables seamless interactions across tasks at different granularities and brings mutual benefits by learning a common and rich pixel-level visual-semantic understanding space, without any pseudo-labeling. After pretraining on a mixed set of a limited amount of segmentation data and millions of image-text pairs, X-Decoder exhibits strong transferability to a wide range of downstream tasks in both zero-shot and finetuning settings. Notably, it achieves (1) state-of-the-art results on open-vocabulary segmentation and referring segmentation on eight datasets; (2) better or competitive finetuned performance to other generalist and specialist models on segmen
中文速览
像素级视觉理解与自然语言理解长期由各自专用模型分别处理,难以相互促进。X-Decoder提出了一种统一解码框架,通过两类查询——不含语义信息的潜在查询和由文字输入生成的语义查询——在同一个语义空间中同时预测像素级分割掩码和语言标记,从而将通用分割、指代分割、图文检索、图像描述生成、视觉问答等任务全部纳入一个模型。模型在少量分割标注数据与海量图文对上端到端预训练,无需借助伪标签,且图像编码器与文本编码器完全解耦,使对比学习与生成学习可同时进行。实验表明,X-Decoder在七个数据集的十项开放词汇分割与指代分割设置上达到最优水平,在多个视觉语言任务上也与专用模型持平甚至更优,还能灵活支持指代式描述生成、图像编辑等新颖任务组合。这项工作首次打通了像素级与图像级视觉语言理解,为构建真正通用的视觉感知基础模型提供了切实可行的路径。
原文 arXiv:2212.11270;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2212.11270v1