Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Li, Rui; Pu, Dong; Huang, Minnie; Huang, Bill

Computer Science > Sound

arXiv:2109.11115 (cs)

[Submitted on 23 Sep 2021 (v1), last revised 24 Feb 2022 (this version, v3)]

Title:Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Authors:Rui Li, Dong Pu, Minnie Huang, Bill Huang

View PDF

Abstract:One-shot voice cloning aims to transform speaker voice and speaking style in speech synthesized from a text-to-speech (TTS) system, where only a shot recording from the target reference speech can be used. Out-of-domain transfer is still a challenging task, and one important aspect that impacts the accuracy and similarity of synthetic speech is the conditional representations carrying speaker or style cues extracted from the limited references. In this paper, we present a novel one-shot voice cloning algorithm called Unet-TTS that has good generalization ability for unseen speakers and styles. Based on a skip-connected U-net structure, the new model can efficiently discover speaker-level and utterance-level spectral feature details from the reference audio, enabling accurate inference of complex acoustic characteristics as well as imitation of speaking styles into the synthetic speech. According to both subjective and objective evaluations of similarity, the new model outperforms both speaker embedding and unsupervised style modeling (GST) approaches on an unseen emotional corpus.

Comments:	6 pages, 5 figures, Accepted to IEEE ICASSP 2022
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2109.11115 [cs.SD]
	(or arXiv:2109.11115v3 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2109.11115

Submission history

From: Rui Li [view email]
[v1] Thu, 23 Sep 2021 03:04:34 UTC (1,208 KB)
[v2] Wed, 29 Sep 2021 06:57:57 UTC (1,209 KB)
[v3] Thu, 24 Feb 2022 10:33:30 UTC (1,212 KB)

Computer Science > Sound

Title:Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators