arXiv:1511.04561 · 中英对照阅读
深度学习中用于并行计算的8位近似
8-Bit Approximations for Parallelism in Deep Learning
中文速览
大规模深度学习并行训练常被 GPU 之间的通信带宽拖慢,导致增加设备后速度提升有限。研究者设计了把 32 位梯度和非线性激活值压缩成 8 位近似值的方法,并结合数据并行和模型并行,在 MNIST、CIFAR10 和 ImageNet 上验证其效果。结果显示,8 位压缩几乎不损害预测准确率,数据传输速度约提升 2 倍,在 96 块 GPU 上预计可获得 50 倍以上加速,而 32 位方法最多约 23 倍。它的重要性在于用较小的精度代价缓解了大规模 GPU 集群的通信瓶颈,让卷积网络更容易扩展到更多设备。
摘要
The creation of practical deep learning data-products often requires parallelization across processors and computers to make deep learning feasible on large data sets, but bottlenecks in communication bandwidth make it difficult to attain good speedups through parallelism. Here we develop and test 8-bit approximation algorithms which make better use of the available bandwidth by compressing 32-bit gradients and nonlinear activations to 8-bit approximations. We show that these approximations do not decrease predictive performance on MNIST, CIFAR10, and ImageNet for both model and data parallelism and provide a data transfer speedup of 2x relative to 32-bit parallelism. We build a predictive model for speedups based on our experimental data, verify its validity on known speedup data, and show that we can obtain a speedup of 50x and more on a system of 96 GPUs compared to a speedup of 23x for 32-bit. We compare our data types with other methods and show that 8-bit approximations achieve state-of-the-art speedups for model parallelism. Thus 8-bit approximation is an efficient method to parallelize convolutional networks on very large systems of GPUs.
术语表
- deep learning
- 深度学习
- data-product
- 数据产品
- parallelization
- 并行化
- parallelism
- 并行计算
- communication bandwidth
- 通信带宽
- speedup
- 加速比
- 8-bit approximation
- 8位近似
- 32-bit gradient
- 32位梯度
- nonlinear activation
- 非线性激活
- model parallelism
- 模型并行
- data parallelism
- 数据并行
- predictive performance
- 预测性能
- MNIST
- MNIST
- CIFAR10
- CIFAR10
- ImageNet
- ImageNet
- stochastic gradient descent (SGD)
- 随机梯度下降(SGD)
- backpropagation
- 反向传播
- gradient
- 梯度
- parameter update
- 参数更新
- mini-batch
- 小批量
- convolutional network
- 卷积网络
- convolutional layer
- 卷积层
- fully connected layer
- 全连接层
- GPU cluster
- GPU集群
- Graphics Processing Unit (GPU)
- 图形处理器(GPU)
- PCI Express (PCIe)
- PCI Express(PCIe)
- InfiniBand
- InfiniBand
- collective communication
- 集合通信
- exploding gradient problem
- 梯度爆炸问题