AdaSpeech: Adaptive Text to Speech for Custom Voice
Mingjian Chen Xu Tan Thanks: The first two authors contribute equally to this work. Corresponding author: Xu Tan, Bohan Li Yanqing Liu Tao Qin Sheng Zhao Tie-Yan LiuMicrosoft Research Asia, Microsoft Azure
Abstract
Custom voice, a specific text to speech (TTS) service in commercial speech platforms, aims to adapt a source TTS model to synthesize personal voice for a target speaker using few speech from her/him. Custom voice presents two unique challenges for TTS adaptation: 1) to support diverse customers, the adaptation model needs to handle diverse acoustic conditions which could be very different from source speech data, and 2) to support a large number of customers, the adaptation parameters need to be small enough for each target speaker to reduce memory usage while maintaining high voice quality. In this work, we propose AdaSpeech, an adaptive TTS system for high-quality and efficient customization of new voices. We design several techniques in AdaSpeech to address the two challenges in custom voice: 1) To handle different acoustic conditions, we model the acoustic information in both utterance and phoneme level. Specifically, we use one acoustic encoder to extract an utterance-level vector and another one to extract a sequence of phoneme-level vectors from the target speech during pre-training and fine-tuning; in inference, we extract the utterance-level vector from a reference speech
原文 arXiv:2103.00993;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2103.00993v1