Mixtral of Experts
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, William El Sayed
Abstract
We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks (i.e. experts). For every token, at each layer, a router network selects two experts to process the current state and combine their outputs. Even though each token only sees two experts, the selected experts can be different at each timestep. As a result, each token has access to 47B parameters, but only uses 13B active parameters during inference. Mixtral was trained with a context size of 32k tokens and it outperforms or matches Llama 2 70B and GPT-3.5 across all evaluated benchmarks. In particular, Mixtral vastly outperforms Llama 2 70B on mathematics, code generation, and multilingual benchmarks. We also provide a model fine-tuned to follow instructions, Mixtral 8x7B – Instruct, that surpasses GPT-3.5 Turbo, Claude-2.1, Gemini Pro, and Llama 2 70B – chat model on human benchmarks. Both the base and instruct models are released under the Apache 2.0 license.
中文速览
Mixtral 8x7B 是一个稀疏混合专家(Sparse Mixture of Experts,SMoE)语言模型,它要解决的核心问题是:如何在不大幅增加推理计算量的前提下,大幅提升模型的参数规模和性能。具体做法是把每一层的前馈网络替换成8个"专家"子网络,每次处理一个token时,由一个路由网络动态选出其中最合适的2个专家来计算,这样模型虽然总共拥有约47B参数,但每个token实际只激活约13B参数,推理成本远低于同量级的稠密模型。实验结果表明,Mixtral 8x7B 在数学、代码生成、多语言理解等多项基准上超越了参数量更大的 Llama 2 70B,并与 GPT-3.5 持平甚至更优;经过指令微调的 Mixtral 8x7B Instruct 版本更在人类评测上击败了 GPT-3.5 Turbo、Claude-2.1 和 Gemini Pro。这项工作的重要意义在于,它以完全开源(Apache 2.0)的方式证明了 SMoE 架构能以更低的推理代价达到顶级模型的性能水平,为学术界和工业界提供了一个高效、可商用的强力基座模型。
原文 arXiv:2401.04088;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2401.04088v1