Scalable and Efficient MoE Training for Multitask Multilingual Models
Young Jin Kim Ammar Ahmad AwanAlexandre Muzio, Andres Felipe Cruz Salinas, Liyang Lu, Amr Hendy Samyam Rajbhandari, Yuxiong He, Hany Hassan AwadallaMicrosoft, One Microsoft Way, Redmond, WA 98052,
Abstract
The Mixture of Experts (MoE) models are an emerging class of sparsely activated deep learning models that have sublinear compute costs with respect to their parameters. In contrast with dense models, the sparse architecture of MoE offers opportunities for drastically growing model size with significant accuracy gain while consuming much lower compute budget. However, supporting large scale MoE training also has its own set of system and modeling challenges.
原文 arXiv:2109.10465;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2109.10465v1