Convolutional Recurrent Neural Network Based Progressive Learning for Monaural Speech Enhancement

Li, Andong; Yuan, Minmin; Zheng, Chengshi; Li, Xiaodong

Computer Science > Sound

arXiv:1908.10768 (cs)

[Submitted on 28 Aug 2019 (v1), last revised 11 Jan 2020 (this version, v3)]

Title:Convolutional Recurrent Neural Network Based Progressive Learning for Monaural Speech Enhancement

Authors:Andong Li, Minmin Yuan, Chengshi Zheng, Xiaodong Li

View PDF

Abstract:Recently, progressive learning has shown its capacity to improve speech quality and speech intelligibility when it is combined with deep neural network (DNN) and long short-term memory (LSTM) based monaural speech enhancement algorithms, especially in low signal-to-noise ratio (SNR) conditions. Nevertheless, due to a large number of parameters and high computational complexity, it is hard to implement in current resource-limited micro-controllers and thus, it is essential to significantly reduce both the number of parameters and the computational load for practical applications. For this purpose, we propose a novel progressive learning framework with causal convolutional recurrent neural networks called PL-CRNN, which takes advantage of both convolutional neural networks and recurrent neural networks to drastically reduce the number of parameters and simultaneously improve speech quality and speech intelligibility. Numerous experiments verify the effectiveness of the proposed PL-CRNN model and indicate that it yields consistent better performance than the PL-DNN and PL-LSTM algorithms and also it gets results close even better than the CRNN in terms of objective measurements. Compared with PL-DNN, PL-LSTM, and CRNN, the proposed PL-CRNN algorithm can reduce the number of parameters up to 93%, 97%, and 92%, respectively.

Comments:	23 pages,5 figures, Submitted to Applied Acoustics
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:1908.10768 [cs.SD]
	(or arXiv:1908.10768v3 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.1908.10768

Submission history

From: Andong Li [view email]
[v1] Wed, 28 Aug 2019 15:09:39 UTC (3,687 KB)
[v2] Tue, 7 Jan 2020 07:07:15 UTC (2,671 KB)
[v3] Sat, 11 Jan 2020 14:06:29 UTC (2,820 KB)

Computer Science > Sound

Title:Convolutional Recurrent Neural Network Based Progressive Learning for Monaural Speech Enhancement

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Convolutional Recurrent Neural Network Based Progressive Learning for Monaural Speech Enhancement

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators