DetNet: A Backbone network for Object Detection
Zeming Li1 Chao Peng2 Gang Yu2 Xiangyu Zhang2 Yangdong Deng1 Jian Sun2 1School of Software Tsinghua University } 2 Megvii Inc. (Face++) {pengchao yugang zhangxiangyu
Abstract
Recent CNN based object detectors, no matter one-stage methods like YOLO [1, 2], SSD [3], and RetinaNet [4] or two-stage detectors like Faster R-CNN [5], R-FCN [6] and FPN [7] are usually trying to directly finetune from ImageNet pre-trained models designed for image classification. There has been little work discussing on the backbone feature extractor specifically designed for the object detection. More importantly, there are several differences between the tasks of image classification and object detection. (i) Recent object detectors like FPN and RetinaNet usually involve extra stages against the task of image classification to handle the objects with various scales. (ii) Object detection not only needs to recognize the category of the object instances but also spatially locate the position. Large downsampling factor brings large valid receptive field, which is good for image classification but compromises the object location ability. Due to the gap between the image classification and object detection, we propose DetNet in this paper, which is a novel backbone network specifically designed for object detection. Moreover, DetNet includes the extra stages against traditional bac
原文 arXiv:1804.06215;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1804.06215v2