MLlib: Machine Learning in Apache Spark
Xiangrui Meng 160 Spear Street, 13th Floor, San Francisco, CA 94105Joseph Bradley 160 Spear Street, 13th Floor, San Francisco, CA 94105Burak Yavuz 160 Spear Street, 13th Floor, San Francisco, CA 94105Evan Sparks Berkeley, 465 Soda Hall, Berkeley, CA 94720Shivaram Venkataraman Berkeley, 465 Soda Hall, Berkeley, CA 94720Davies Liu 160 Spear Street, 13th Floor, San Francisco, CA 94105Jeremy Freeman Janelia Research Campus, 19805 Helix Dr, Ashburn, VA 20147DB Tsai 970 University Ave, Los Gatos, CA 95032Manish Amde Logic, 1134 Crane Street, Menlo Park, CA 94025Sean Owen UK, 33 Creechurch Lane, London EC3A 5EB United KingdomDoris Xin 201 N Goodwin Ave, Urbana, IL 61801Reynold 160 Spear Street, 13th Floor, San Francisco, CA 94105Michael J. Franklin Berkeley, 465 Soda Hall, Berkeley, CA 94720Reza Zadeh and Databricks, 475 Via Ortega, Stanford, CA 94305Matei Zaharia and Databricks, 160 Spear Street, 13th Floor, San Francisco, CA 94105 Ameet Talwalkar and Databricks, 4732 Boelter Hall, Los Angeles, CA 90095
Abstract
Apache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark’s open-source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes several underlying statistical, optimization, and linear algebra primitives. Shipped with Spark, MLlib supports several languages and provides a high-level API that leverages Spark’s rich ecosystem to simplify the development of end-to-end machine learning pipelines. MLlib has experienced a rapid growth due to its vibrant open-source community of over 140 contributors, and includes extensive documentation to support further growth and to let users quickly get up to speed.
原文 arXiv:1505.06807;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1505.06807v1