Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling https://github.com/tensorflow/lingvo
Jonathan Shen Patrick Nguyen Yonghui Wu Zhifeng Chen Mia X. Chen Ye Jia Anjuli Kannan Tara Sainath Yuan Cao Chung-Cheng Chiu Yanzhang He Jan Chorowski Smit Hinsu Stella Laurenzo James Qin Orhan Firat Wolfgang Macherey Suyog Gupta Ankur Bapna Shuyuan Zhang Ruoming Pang Ron J. Weiss Rohit Prabhavalkar Qiao Liang Benoit Jacob Bowen Liang HyoukJoong Lee Ciprian Chelba Sébastien Jean Bo Li Melvin Johnson Rohan Anil Rajat Tibrewal Xiaobing Liu Akiko Eriguchi Navdeep Jaitly Naveen Ari Colin Cherry Parisa Haghani Otavio Good Youlong Cheng Raziel Alvarez Isaac Caswell Wei-Ning Hsu Zongheng Yang Kuan-Chieh Wang Ekaterina Gonina Katrin Tomanek Ben Vanik Zelin Wu Llion Jones Mike Schuster Yanping Huang Dehao Chen Kazuki Irie George Foster John Richardson Klaus Macherey Antoine Bruguier Heiga Zen Colin Raffel Shankar Kumar Kanishka Rao David Rybach Matthew Murray Vijayaditya Peddinti Maxim Krikun Michiel A. U. Bacchiani Thomas B. Jablin Rob Suderman Ian Williams Benjamin Lee Deepti Bhatia Justin Carlson Semih Yavuz Yu Zhang Ian McGraw Max Galkin Qi Ge Golan Pundak Chad Whipkey Todd Wang Uri Alon Dmitry Lepikhin Ye Tian Sara Sabour William Chan Shubham Toshniwal Baohua Liao Michael Nirschl Pat Rondon
Abstract
Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, and experiment configurations are centralized and highly customizable. Distributed training and quantized inference are supported directly within the framework, and it contains existing implementations of a large number of utilities, helper functions, and the newest research ideas. Lingvo has been used in collaboration by dozens of researchers in more than 20 papers over the last two years. This document outlines the underlying design of Lingvo and serves as an introduction to the various pieces of the framework, while also offering examples of advanced features that showcase the capabilities of the framework.
中文速览
序列建模研究中往往面临代码难以复用、实验难以复现、从研究到部署需大量改写等痛点,谷歌为此开发了开源深度学习框架 Lingvo,并在这篇文章中系统介绍了它的设计思路与核心组件。Lingvo 以 TensorFlow 为基础,将模型拆分为层(Layer)、任务(Task)、模型(Model)等模块化积木,所有超参数集中在独立的配置类中管理,训练、评估、推理由不同的 Job Runner 分工执行,并原生支持同步/异步分布式训练和量化推理。得益于这套设计,过去两年间数十位研究者共用同一代码库,在机器翻译、语音识别、语音合成、语音翻译等领域发表了逾二十篇论文并取得多项最优结果。这项工作的价值在于它提供了一套兼顾快速迭代与工程规范的研究基础设施,让算法改进能在不同任务间直接复用,同时使实验可共享、可比较、可复现,对需要大规模协作的深度学习研究具有重要参考意义。
原文 arXiv:1902.08295;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1902.08295v1