Learning Holistic Geometric Representations for Monocular 3D Object Detection
Yinmin Zhang1, Xinzhu Ma2, Shuai Yi1, Jun Hou1, Zhihui Wang3, Wanli Ouyang2 Dan Xu4, 1SenseTime Research, 2The University of Sydney, 3Dalian University of Technology 4The Hong Kong University of Science and Technology, {zhangyinmin, yishuai, {xinzhu.ma,
Abstract
As a crucial task of autonomous driving, 3D object detection has made significant progress in recent years. However, monocular 3D object detection remains a challenging problem due to the unsatisfactory performance in depth estimation. Most existing monocular methods typically directly regress the depth, while ignoring essential relationships between the depth and various geometric elements (e.g. bounding box sizes, 3D object dimensions, and object poses). In this paper, we propose to learn geometry-guided depth estimation with projective modeling to advance monocular 3D object detection. Specifically, a principled geometry formula with projective modeling of 2D and 3D depth predictions in the monocular 3D object detection network is devised. We further implement and embed the proposed formula to enable geometry-aware deep representation learning, allowing effective 2D and 3D interactions for boosting the depth estimation. Moreover, we provide a strong baseline through addressing substantial misalignment between 2D annotation and projected boxes to ensure robust learning with the proposed holistic geometric formula. Experiments on the KITTI dataset show that our method remarkably i
原文 arXiv:2107.13931;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2107.13931v2