Training Text-To-Speech Systems From Synthetic Data: A Practical Approach For Accent Transfer Tasks

Finkelstein, Lev; Zen, Heiga; Casagrande, Norman; Chan, Chun-an; Jia, Ye; Kenter, Tom; Petelin, Alexey; Shen, Jonathan; Wan, Vincent; Zhang, Yu; Wu, Yonghui; Clark, Rob

Computer Science > Sound

arXiv:2208.13183 (cs)

[Submitted on 28 Aug 2022]

Title:Training Text-To-Speech Systems From Synthetic Data: A Practical Approach For Accent Transfer Tasks

Authors:Lev Finkelstein, Heiga Zen, Norman Casagrande, Chun-an Chan, Ye Jia, Tom Kenter, Alexey Petelin, Jonathan Shen, Vincent Wan, Yu Zhang, Yonghui Wu, Rob Clark

View PDF

Abstract:Transfer tasks in text-to-speech (TTS) synthesis - where one or more aspects of the speech of one set of speakers is transferred to another set of speakers that do not feature these aspects originally - remains a challenging task. One of the challenges is that models that have high-quality transfer capabilities can have issues in stability, making them impractical for user-facing critical tasks. This paper demonstrates that transfer can be obtained by training a robust TTS system on data generated by a less robust TTS system designed for a high-quality transfer task; in particular, a CHiVE-BERT monolingual TTS system is trained on the output of a Tacotron model designed for accent transfer. While some quality loss is inevitable with this approach, experimental results show that the models trained on synthetic data this way can produce high quality audio displaying accent transfer, while preserving speaker characteristics such as speaking style.

Comments:	To be published in Interspeech 2022
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2208.13183 [cs.SD]
	(or arXiv:2208.13183v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2208.13183

Submission history

From: Lev Finkelstein [view email]
[v1] Sun, 28 Aug 2022 09:24:45 UTC (149 KB)

Computer Science > Sound

Title:Training Text-To-Speech Systems From Synthetic Data: A Practical Approach For Accent Transfer Tasks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Training Text-To-Speech Systems From Synthetic Data: A Practical Approach For Accent Transfer Tasks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators