arXiv:1511.06051 · 中英对照阅读
SparkNet: Training Deep Networks in Spark
中文速览
深度网络训练往往要耗费数天,而常用的 Spark 等批处理框架又不适合频繁通信、异步更新的分布式训练。SparkNet 把 Spark 与 Caffe 结合起来,让各机器在本地连续进行一段时间的随机梯度下降(SGD),再定期汇总并平均模型参数,从而大幅减少通信,同时兼容现有 Caffe 模型和 Spark 数据处理流程。实验表明,它能随机器数量增加获得良好加速,在通信延迟很高、带宽有限的集群上也能稳定工作,并在 ImageNet 上取得接近专用分布式系统的训练表现。它的重要性在于无需复杂调参或专门硬件,就能把深度学习训练直接嵌入已有的大规模数据处理流水线。
摘要
Training deep networks is a time-consuming process, with networks for object recognition often requiring multiple days to train. For this reason, leveraging the resources of a cluster to speed up training is an important area of work. However, widely-popular batch-processing computational frameworks like MapReduce and Spark were not designed to support the asynchronous and communication-intensive workloads of existing distributed deep learning systems. We introduce SparkNet, a framework for training deep networks in Spark. Our implementation includes a convenient interface for reading data from Spark RDDs, a Scala interface to the Caffe deep learning framework, and a lightweight multi-dimensional tensor library. Using a simple parallelization scheme for stochastic gradient descent, SparkNet scales well with the cluster size and tolerates very high-latency communication. Furthermore, it is easy to deploy and use with no parameter tuning, and it is compatible with existing Caffe models. We quantify the dependence of the speedup obtained by SparkNet on the number of machines, the communication frequency, and the cluster’s communication overhead, and we benchmark our system’s performance on the ImageNet dataset.
术语表
- SparkNet
- SparkNet
- deep learning
- 深度学习
- distributed deep learning
- 分布式深度学习
- MapReduce
- MapReduce
- Spark
- Spark
- batch-processing framework
- 批处理计算框架
- asynchronous, lock-free optimization
- 异步无锁优化
- parameter server model
- 参数服务器模型
- master node
- 主节点
- worker node
- 工作节点
- model parameters
- 模型参数
- stochastic gradient descent (SGD)
- 随机梯度下降(SGD)
- gradient
- 梯度
- minibatch
- 小批量
- parallelization scheme
- 并行化方案
- communication overhead
- 通信开销
- bandwidth-limited environment
- 带宽受限环境
- Spark RDD
- Spark RDD
- Caffe
- Caffe
- Scala
- Scala
- Java Native Access
- Java Native Access
- Google Protocol Buffers
- Google Protocol Buffers
- NDArray
- NDArray
- multi-dimensional tensor library
- 多维张量库
- NetParams
- NetParams
- WeightCollection
- WeightCollection
- ImageNet
- ImageNet
- AlexNet
- AlexNet
- EC2
- EC2