Guided-TTS 2: A Diffusion Model for High-quality Adaptive Text-to-Speech with Untranscribed Data
Sungwon Kim Data Science、AI Lab. Seoul National University、Heeseung Kim Data Science、AI Lab. Seoul National University、Sungroh Yoon Data Science、AI Lab. Seoul National University Equal contribution.Corresponding author
Abstract
We propose Guided-TTS 2, a diffusion-based generative model for high-quality adaptive TTS using untranscribed data. Guided-TTS 2 combines a speaker-conditional diffusion model with a speaker-dependent phoneme classifier for adaptive text-to-speech. We train the speaker-conditional diffusion model on large-scale untranscribed datasets for a classifier-free guidance method and further fine-tune the diffusion model on the reference speech of the target speaker for adaptation, which only takes 40 seconds. We demonstrate that Guided-TTS 2 shows comparable performance to high-quality single-speaker TTS baselines in terms of speech quality and speaker similarity with only a ten-second untranscribed data. We further show that Guided-TTS 2 outperforms adaptive TTS baselines on multi-speaker datasets even with a zero-shot adaptation setting. Guided-TTS 2 can adapt to a wide range of voices only using untranscribed speech, which enables adaptive TTS with the voice of non-human characters such as Gollum in "The Lord of the Rings".
中文速览
让语音合成系统用极少量数据就能逼真模仿任意说话人的声音,这是自适应语音合成(adaptive TTS)领域的核心难题。Guided-TTS 2 的做法是:先用海量无标注(无文字转录)多说话人语音训练一个以说话人身份为条件的扩散模型(speaker-conditional diffusion model),再仅用目标说话人约10秒的无标注语音对该模型进行快速微调(整个适配过程不到40秒),同时借助预训练音素分类器引导扩散模型的采样过程以保证发音准确性。实验结果显示,用10秒LJSpeech语音微调后,其语音质量和说话人相似度已接近用24小时数据训练的单说话人高质量TTS系统,并在多说话人基准测试中超越了现有自适应TTS方法;更令人惊喜的是,由于整个流程完全不依赖文字转录,该方法甚至能成功模仿《指环王》中咕噜这类难以转录的非人类角色声音,极大拓展了自适应语音合成的应用边界。
原文 arXiv:2205.15370;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.15370v1