Aha.
正在载入中英对照阅读…

arXiv:1511.04561 · 中英对照阅读

深度学习中用于并行计算的8位近似

8-Bit Approximations for Parallelism in Deep Learning

Tim Dettmers

中文速览

大规模深度学习并行训练常被 GPU 之间的通信带宽拖慢,导致增加设备后速度提升有限。研究者设计了把 32 位梯度和非线性激活值压缩成 8 位近似值的方法,并结合数据并行和模型并行,在 MNIST、CIFAR10 和 ImageNet 上验证其效果。结果显示,8 位压缩几乎不损害预测准确率,数据传输速度约提升 2 倍,在 96 块 GPU 上预计可获得 50 倍以上加速,而 32 位方法最多约 23 倍。它的重要性在于用较小的精度代价缓解了大规模 GPU 集群的通信瓶颈,让卷积网络更容易扩展到更多设备。

摘要

The creation of practical deep learning data-products often requires parallelization across processors and computers to make deep learning feasible on large data sets, but bottlenecks in communication bandwidth make it difficult to attain good speedups through parallelism. Here we develop and test 8-bit approximation algorithms which make better use of the available bandwidth by compressing 32-bit gradients and nonlinear activations to 8-bit approximations. We show that these approximations do not decrease predictive performance on MNIST, CIFAR10, and ImageNet for both model and data parallelism and provide a data transfer speedup of 2x relative to 32-bit parallelism. We build a predictive model for speedups based on our experimental data, verify its validity on known speedup data, and show that we can obtain a speedup of 50x and more on a system of 96 GPUs compared to a speedup of 23x for 32-bit. We compare our data types with other methods and show that 8-bit approximations achieve state-of-the-art speedups for model parallelism. Thus 8-bit approximation is an efficient method to parallelize convolutional networks on very large systems of GPUs.

术语表

deep learning
深度学习
data-product
数据产品
parallelization
并行化
parallelism
并行计算
communication bandwidth
通信带宽
speedup
加速比
8-bit approximation
8位近似
32-bit gradient
32位梯度
nonlinear activation
非线性激活
model parallelism
模型并行
data parallelism
数据并行
predictive performance
预测性能
MNIST
MNIST
CIFAR10
CIFAR10
ImageNet
ImageNet
stochastic gradient descent (SGD)
随机梯度下降(SGD)
backpropagation
反向传播
gradient
梯度
parameter update
参数更新
mini-batch
小批量
convolutional network
卷积网络
convolutional layer
卷积层
fully connected layer
全连接层
GPU cluster
GPU集群
Graphics Processing Unit (GPU)
图形处理器(GPU)
PCI Express (PCIe)
PCI Express(PCIe)
InfiniBand
InfiniBand
collective communication
集合通信
exploding gradient problem
梯度爆炸问题