MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

Yang, Qian; Zuo, Jialong; Su, Zhe; Jiang, Ziyue; Li, Mingze; Zhao, Zhou; Chen, Feiyang; Wang, Zhefeng; Huai, Baoxing

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2407.14006 (eess)

[Submitted on 19 Jul 2024]

Title:MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

Authors:Qian Yang, Jialong Zuo, Zhe Su, Ziyue Jiang, Mingze Li, Zhou Zhao, Feiyang Chen, Zhefeng Wang, Baoxing Huai

View PDF HTML (experimental)

Abstract:We introduce an open source high-quality Mandarin TTS dataset MSceneSpeech (Multiple Scene Speech Dataset), which is intended to provide resources for expressive speech synthesis. MSceneSpeech comprises numerous audio recordings and texts performed and recorded according to daily life scenarios. Each scenario includes multiple speakers and a diverse range of prosodic styles, making it suitable for speech synthesis that entails multi-speaker style and prosody modeling. We have established a robust baseline, through the prompting mechanism, that can effectively synthesize speech characterized by both user-specific timbre and scene-specific prosody with arbitrary text input. The open source MSceneSpeech Dataset and audio samples of our baseline are available at this https URL.

Comments:	Accepted by INTERSPEECH 2024
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2407.14006 [eess.AS]
	(or arXiv:2407.14006v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2407.14006

Submission history

From: Qian Yang [view email]
[v1] Fri, 19 Jul 2024 03:36:48 UTC (326 KB)

Full-text links:

Access Paper:

view license

Current browse context:

eess.AS

< prev | next >

new | recent | 2024-07

Change to browse by:

cs
cs.SD
eess

References & Citations

export BibTeX citation

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators