arXiv:1802.08530 · 中英对照阅读
Training wide residual networks for deployment using a single bit for each weight
中文速览
论文要解决的是:如何把深度卷积网络的每个权重压到1 bit,从而让模型更省内存、更省电、适合嵌入式设备,同时尽量不牺牲识别准确率。作者以宽残差网络为基础,训练时让权重在前向和反向传播中取正负1,但仍用全精度权重更新,并为每层加入根据初始化标准差设定的固定缩放因子,同时在容易过拟合的任务中固定批归一化参数、配合带热重启的学习率策略。结果显示,在CIFAR-10、CIFAR-100和ImageNet上,1 bit权重模型分别达到3.9%、18.5%和26.0%/8.5%的错误率,CIFAR数据上的错误率约为此前方法的一半,并且与全精度模型只差约1%。这说明极低精度权重不必以明显准确率下降为代价,经过简单训练改动就能得到接近全精度的模型,为低功耗芯片、传感器、机器人和物联网设备部署深度网络提供了更实用的路径。
摘要
For fast and energy-efficient deployment of trained deep neural networks on resource-constrained embedded hardware, each learned weight parameter should ideally be represented and stored using a single bit. Error-rates usually increase when this requirement is imposed. Here, we report large improvements in error rates on multiple datasets, for deep convolutional neural networks deployed with 1-bit-per-weight. Using wide residual networks as our main baseline, our approach simplifies existing methods that binarize weights by applying the sign function in training; we apply scaling factors for each layer with constant unlearned values equal to the layer-specific standard deviations used for initialization. For CIFAR-10, CIFAR-100 and ImageNet, and models with 1-bit-per-weight requiring less than 10 MB of parameter memory, we achieve error rates of 3.9%, 18.5% and 26.0% / 8.5% (Top-1 / Top-5) respectively. We also considered MNIST, SVHN and ImageNet32, achieving 1-bit-per-weight test results of 0.27%, 1.9%, and 41.3% / 19.1% respectively. For CIFAR, our error rates halve previously reported values, and are within about 1% of our error-rates for the same network with full-precision weights. For networks that overfit, we also show significant improvements in error rate by not learning batch normalization scale and offset parameters. This applies to both full precision and 1-bit-per-weight networks. Using a warm-restart learning-rate schedule, we found that training for 1-bit-per-weight is just as fast as full-precision networks, with better accuracy than standard schedules, and achieved about 98%-99% of peak performance in just 62 training epochs for CIFAR-10/100. For full training code and trained models in MATLAB, Keras and PyTorch see https://github.com/McDonnell-Lab/1-bit-per-weight/.
术语表
- 1-bit-per-weight
- 每权重 1 位
- deep neural network (DNN)
- 深度神经网络(DNN)
- resource-constrained embedded hardware
- 资源受限嵌入式硬件
- weight binarization
- 权重二值化
- sign function
- 符号函数
- scaling factor
- 缩放因子
- standard deviation
- 标准差
- convolutional neural network (CNN)
- 卷积神经网络(CNN)
- Wide Residual Network (WRN)
- 宽残差网络(Wide Residual Network,WRN)
- ResNet
- 残差网络(ResNet)
- full-precision weights
- 全精度权重
- single-precision floating point
- 单精度浮点数
- batch normalization
- 批归一化
- batch normalization scale and offset parameters
- 批归一化缩放与偏移参数
- warm-restart learning-rate schedule
- 热重启学习率调度
- Top-1 accuracy
- Top-1 准确率
- Top-5 accuracy
- Top-5 准确率
- global-average-pooling
- 全局平均池化
- skip-connections
- 跳跃连接
- FLOPs (FLoating-point OPerations)
- 浮点运算次数(FLOPs)
- model compression
- 模型压缩
- weight pruning
- 权重剪枝
- weight quantization
- 权重量化
- forward propagation
- 前向传播
- inference
- 推理
- CIFAR-10
- CIFAR-10
- CIFAR-100
- CIFAR-100
- ImageNet
- ImageNet
- MNIST
- MNIST
- SVHN
- SVHN