Accelerating Convolutional Neural Network by Exploiting Sparsity on GPUsDOI: XXXXXXX.XXXXXXXNote: New Paper, Not an Extension of a Conference PaperCCS: Computing methodologies Parallel algorithmsCCS: Computing methodologies Machine learning
Weizhi Xu Affiliation: School of Information Science and Engineering, Shandong Normal University , Jinan , China Affiliation: Dept. of Electrical、Computer Engineering, University of Houston , Houston , USA , Yintai Sun Affiliation: School of Information Science and Engineering, Shandong Normal University , Jinan , China , Shengyu Fan Affiliation: School of Information Science and Engineering, Shandong Normal University , Jinan , China , Hui Yu Affiliation: School of Information Science and Engineering, Shandong Normal University , Jinan , China Affiliation: Dept. of Electrical、Computer Engineering, University of Houston , Houston , USA and Xin Fu Affiliation: Dept. of Electrical、Computer Engineering, University of Houston , Houston , USA Note: Corresponding author
Abstract
Convolution Neural Network (CNN) is an important deep learning method, which is widely used in many fields. However, it is very time consuming to implement CNN where convolution usually takes most of the time. There are many zero values in feature maps and filters, which leads to redundant calculations and memory accesses if dense methods are used to compute convolution. Many works recently make use of sparsity to skip the calculations for zero values to reduce the inference time of CNN. On the GPU platform, current works cannot fully exploit the sparsity of the feature map and achieve satisfactory performance. Therefore, we design a new parallel strategy to transform the feature map into a new storage format to avoid the redundant computation of zero-values on GPUs. Also considering the sparsity in the feature map, we propose a fused storage format to combine the convolution operation with the following pooling operation, in order to further improve the performance. We carry out experiments with mainstream CNN models and achieve better performance compared with cuDNN and cuSPARSE. For VGG-19, ResNet-50, DenseNet-121 and RegNetX-16GF, $1.97\times$ , $2.23\times$ , $2.74\times$ and
原文 arXiv:1909.09927;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1909.09927v6