Aha.
正在载入中英对照阅读…

arXiv:1511.06051 · 中英对照阅读

SparkNet: Training Deep Networks in Spark

Philipp Moritz、Robert Nishihara*、Ion Stoica、Michael I. Jordan

中文速览

深度网络训练往往要耗费数天,而常用的 Spark 等批处理框架又不适合频繁通信、异步更新的分布式训练。SparkNet 把 Spark 与 Caffe 结合起来,让各机器在本地连续进行一段时间的随机梯度下降(SGD),再定期汇总并平均模型参数,从而大幅减少通信,同时兼容现有 Caffe 模型和 Spark 数据处理流程。实验表明,它能随机器数量增加获得良好加速,在通信延迟很高、带宽有限的集群上也能稳定工作,并在 ImageNet 上取得接近专用分布式系统的训练表现。它的重要性在于无需复杂调参或专门硬件,就能把深度学习训练直接嵌入已有的大规模数据处理流水线。

摘要

Training deep networks is a time-consuming process, with networks for object recognition often requiring multiple days to train. For this reason, leveraging the resources of a cluster to speed up training is an important area of work. However, widely-popular batch-processing computational frameworks like MapReduce and Spark were not designed to support the asynchronous and communication-intensive workloads of existing distributed deep learning systems. We introduce SparkNet, a framework for training deep networks in Spark. Our implementation includes a convenient interface for reading data from Spark RDDs, a Scala interface to the Caffe deep learning framework, and a lightweight multi-dimensional tensor library. Using a simple parallelization scheme for stochastic gradient descent, SparkNet scales well with the cluster size and tolerates very high-latency communication. Furthermore, it is easy to deploy and use with no parameter tuning, and it is compatible with existing Caffe models. We quantify the dependence of the speedup obtained by SparkNet on the number of machines, the communication frequency, and the cluster’s communication overhead, and we benchmark our system’s performance on the ImageNet dataset.

术语表

SparkNet
SparkNet
deep learning
深度学习
distributed deep learning
分布式深度学习
MapReduce
MapReduce
Spark
Spark
batch-processing framework
批处理计算框架
asynchronous, lock-free optimization
异步无锁优化
parameter server model
参数服务器模型
master node
主节点
worker node
工作节点
model parameters
模型参数
stochastic gradient descent (SGD)
随机梯度下降(SGD)
gradient
梯度
minibatch
小批量
parallelization scheme
并行化方案
communication overhead
通信开销
bandwidth-limited environment
带宽受限环境
Spark RDD
Spark RDD
Caffe
Caffe
Scala
Scala
Java Native Access
Java Native Access
Google Protocol Buffers
Google Protocol Buffers
NDArray
NDArray
multi-dimensional tensor library
多维张量库
NetParams
NetParams
WeightCollection
WeightCollection
ImageNet
ImageNet
AlexNet
AlexNet
EC2
EC2