Aha.
正在载入中英对照阅读…

arXiv:1802.08530 · 中英对照阅读

Training wide residual networks for deployment using a single bit for each weight

Mark D. McDonnell

中文速览

论文要解决的是:如何把深度卷积网络的每个权重压到1 bit,从而让模型更省内存、更省电、适合嵌入式设备,同时尽量不牺牲识别准确率。作者以宽残差网络为基础,训练时让权重在前向和反向传播中取正负1,但仍用全精度权重更新,并为每层加入根据初始化标准差设定的固定缩放因子,同时在容易过拟合的任务中固定批归一化参数、配合带热重启的学习率策略。结果显示,在CIFAR-10、CIFAR-100和ImageNet上,1 bit权重模型分别达到3.9%、18.5%和26.0%/8.5%的错误率,CIFAR数据上的错误率约为此前方法的一半,并且与全精度模型只差约1%。这说明极低精度权重不必以明显准确率下降为代价,经过简单训练改动就能得到接近全精度的模型,为低功耗芯片、传感器、机器人和物联网设备部署深度网络提供了更实用的路径。

摘要

For fast and energy-efficient deployment of trained deep neural networks on resource-constrained embedded hardware, each learned weight parameter should ideally be represented and stored using a single bit. Error-rates usually increase when this requirement is imposed. Here, we report large improvements in error rates on multiple datasets, for deep convolutional neural networks deployed with 1-bit-per-weight. Using wide residual networks as our main baseline, our approach simplifies existing methods that binarize weights by applying the sign function in training; we apply scaling factors for each layer with constant unlearned values equal to the layer-specific standard deviations used for initialization. For CIFAR-10, CIFAR-100 and ImageNet, and models with 1-bit-per-weight requiring less than 10 MB of parameter memory, we achieve error rates of 3.9%, 18.5% and 26.0% / 8.5% (Top-1 / Top-5) respectively. We also considered MNIST, SVHN and ImageNet32, achieving 1-bit-per-weight test results of 0.27%, 1.9%, and 41.3% / 19.1% respectively. For CIFAR, our error rates halve previously reported values, and are within about 1% of our error-rates for the same network with full-precision weights. For networks that overfit, we also show significant improvements in error rate by not learning batch normalization scale and offset parameters. This applies to both full precision and 1-bit-per-weight networks. Using a warm-restart learning-rate schedule, we found that training for 1-bit-per-weight is just as fast as full-precision networks, with better accuracy than standard schedules, and achieved about 98%-99% of peak performance in just 62 training epochs for CIFAR-10/100. For full training code and trained models in MATLAB, Keras and PyTorch see https://github.com/McDonnell-Lab/1-bit-per-weight/.

术语表

1-bit-per-weight
每权重 1 位
deep neural network (DNN)
深度神经网络(DNN)
resource-constrained embedded hardware
资源受限嵌入式硬件
weight binarization
权重二值化
sign function
符号函数
scaling factor
缩放因子
standard deviation
标准差
convolutional neural network (CNN)
卷积神经网络(CNN)
Wide Residual Network (WRN)
宽残差网络(Wide Residual Network,WRN)
ResNet
残差网络(ResNet)
full-precision weights
全精度权重
single-precision floating point
单精度浮点数
batch normalization
批归一化
batch normalization scale and offset parameters
批归一化缩放与偏移参数
warm-restart learning-rate schedule
热重启学习率调度
Top-1 accuracy
Top-1 准确率
Top-5 accuracy
Top-5 准确率
global-average-pooling
全局平均池化
skip-connections
跳跃连接
FLOPs (FLoating-point OPerations)
浮点运算次数(FLOPs)
model compression
模型压缩
weight pruning
权重剪枝
weight quantization
权重量化
forward propagation
前向传播
inference
推理
CIFAR-10
CIFAR-10
CIFAR-100
CIFAR-100
ImageNet
ImageNet
MNIST
MNIST
SVHN
SVHN