Natural TTS Synthesis By Conditioning WaveNet On Mel Spectrogram Predictions
Jonathan Shen Google, Inc. Ruoming Pang Google, Inc. Ron J. Weiss Google, Inc. Mike Schuster Google, Inc. Navdeep Jaitly Google, Inc. Zongheng Yang Work done while at Google. University of California, Berkeley Zhifeng Chen Google, Inc. Yu Zhang Google, Inc. Yuxuan Wang Google, Inc. RJ Skerry-Ryan Google, Inc. Rif A. Saurous Google, Inc. Yannis Agiomyrgiannakis Google, Inc. Yonghui Wu Google, Inc.
Abstract
This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by a modified WaveNet model acting as a vocoder to synthesize time-domain waveforms from those spectrograms. Our model achieves a mean opinion score (MOS) of $4.53$ comparable to a MOS of $4.58$ for professionally recorded speech. To validate our design choices, we present ablation studies of key components of our system and evaluate the impact of using mel spectrograms as the conditioning input to WaveNet instead of linguistic, duration, and $F_{0}$ features. We further show that using this compact acoustic intermediate representation allows for a significant reduction in the size of the WaveNet architecture.
中文速览
把文字直接合成出堪比真人录音的语音,长期以来需要繁琐的语言学特征工程,Tacotron 2 用一套端到端神经网络彻底绕开了这一难题:先用一个带注意力机制的序列到序列(sequence-to-sequence)网络,把输入字符直接映射成梅尔频谱(mel spectrogram),再用改进版 WaveNet 把这张"声音地图"还原成真实波形。整套系统在标准主观评分(MOS)测试中拿到 4.53 分,与专业录音的 4.58 分几乎持平,远超此前所有对比系统。这项工作的意义在于,它证明了一张紧凑的梅尔频谱就足以替代复杂的语言学、时长和基频特征,不仅让整个合成流程大幅简化,还让 WaveNet 的模型规模显著缩小,为高质量语音合成的实用化铺平了道路。
原文 arXiv:1712.05884;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1712.05884v2