Self-Supervised Learning based Monaural Speech Enhancement with Multi-Task Pre-Training

Li, Yi; Sun, Yang; Naqvi, Syed Mohsen

Computer Science > Sound

arXiv:2112.11459 (cs)

[Submitted on 21 Dec 2021]

Title:Self-Supervised Learning based Monaural Speech Enhancement with Multi-Task Pre-Training

Authors:Yi Li, Yang Sun, Syed Mohsen Naqvi

View PDF

Abstract:In self-supervised learning, it is challenging to reduce the gap between the enhancement performance on the estimated and target speech signals with existed pre-tasks. In this paper, we propose a multi-task pre-training method to improve the speech enhancement performance with self-supervised learning. Within the pre-training autoencoder (PAE), only a limited set of clean speech signals are required to learn their latent representations. Meanwhile, to solve the limitation of single pre-task, the proposed masking module exploits the dereverberated mask and estimated ratio mask to denoise the mixture as the second pre-task. Different from the PAE, where the target speech signals are estimated, the downstream task autoencoder (DAE) utilizes a large number of unlabeled and unseen reverberant mixtures to generate the estimated mixtures. The trained DAE is shared by the learned representations and masks. Experimental results on a benchmark dataset demonstrate that the proposed method outperforms the state-of-the-art approaches.

Comments:	Submitted to ICASSP 2022. arXiv admin note: text overlap with arXiv:2112.11142
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2112.11459 [cs.SD]
	(or arXiv:2112.11459v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2112.11459

Submission history

From: Yi Li [view email]
[v1] Tue, 21 Dec 2021 13:08:45 UTC (2,141 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.SD

< prev | next >

new | recent | 2021-12

Change to browse by:

cs
eess
eess.AS

References & Citations

DBLP - CS Bibliography

listing | bibtex

Yi Li
Yang Sun
Syed Mohsen Naqvi

export BibTeX citation

Computer Science > Sound

Title:Self-Supervised Learning based Monaural Speech Enhancement with Multi-Task Pre-Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Self-Supervised Learning based Monaural Speech Enhancement with Multi-Task Pre-Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators