FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Yi Ren Thanks: Authors contribute equally to this work. Affiliation: Zhejiang Chenxu Hu Affiliation: Zhejiang Xu Tan Affiliation: Microsoft Research Tao Qin Affiliation: Microsoft Research Sheng Zhao Affiliation: Microsoft Azure Zhou Zhao Thanks: Corresponding author Affiliation: Zhejiang Tie-Yan Liu Affiliation: Microsoft Research
Abstract
Non-autoregressive text to speech (TTS) models such as FastSpeech (Ren et al. 2019) can synthesize speech significantly faster than previous autoregressive models with comparable quality. The training of FastSpeech model relies on an autoregressive teacher model for duration prediction (to provide more information as input) and knowledge distillation (to simplify the data distribution in output), which can ease the one-to-many mapping problem (i.e., multiple speech variations correspond to the same text) in TTS. However, FastSpeech has several disadvantages: 1) the teacher-student distillation pipeline is complicated and time-consuming, 2) the duration extracted from the teacher model is not accurate enough, and the target mel-spectrograms distilled from teacher model suffer from information loss due to data simplification, both of which limit the voice quality. In this paper, we propose FastSpeech 2, which addresses the issues in FastSpeech and better solves the one-to-many mapping problem in TTS by 1) directly training the model with ground-truth target instead of the simplified output from teacher, and 2) introducing more variation information of speech (e.g., pitch, energy and
原文 arXiv:2006.04558;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2006.04558v8