A Simple Framework for Open-Vocabulary Segmentation and Detection
Hao Zhang Thanks: Equal contribution. List in random. Affiliation: The Hong Kong University of Science and Technology. Affiliation: International Digital Economy Academy (IDEA). Feng Li Affiliation: The Hong Kong University of Science and Technology. Xueyan Zou Affiliation: Microsoft Research at Redmond. University of Wisconsin-Madison. Shilong Liu Affiliation: International Digital Economy Academy (IDEA). Affiliation: Dept. of CST., BNRist Center, Institute for AI, Tsinghua University.{hzhangcx, Chunyuan Li Jianfeng Gao Jianwei Yang Thanks: Equal advisory contribution. Lei Zhang
Abstract
We present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets. To bridge the gap of vocabulary and annotation granularity, we first introduce a pre-trained text encoder to encode all the visual concepts in two tasks and learn a common semantic space for them. This gives us reasonably good results compared with the counterparts trained on segmentation task only. To further reconcile them, we locate two discrepancies: $i$ ) task discrepancy – segmentation requires extracting masks for both foreground objects and background stuff, while detection merely cares about the former; $ii$ ) data discrepancy -- box and mask annotations are with different spatial granularity, and thus not directly interchangeable. To address these issues, we propose a decoupled decoding to reduce the interference between foreground/background and a conditioned mask decoding to assist in generating masks for given boxes. To this end, we develop a simple encoder-decoder model encompassing all three techniques and train it jointly on COCO and ††footnotetext: This work is developed during an internship at IDEA. Objects365.
原文 arXiv:2303.08131;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2303.08131v3