End-to-End Speaker-Attributed ASR with Transformer

Kanda, Naoyuki; Ye, Guoli; Gaur, Yashesh; Wang, Xiaofei; Meng, Zhong; Chen, Zhuo; Yoshioka, Takuya

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2104.02128 (eess)

[Submitted on 5 Apr 2021]

Title:End-to-End Speaker-Attributed ASR with Transformer

Authors:Naoyuki Kanda, Guoli Ye, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, Takuya Yoshioka

View PDF

Abstract:This paper presents our recent effort on end-to-end speaker-attributed automatic speech recognition, which jointly performs speaker counting, speech recognition and speaker identification for monaural multi-talker audio. Firstly, we thoroughly update the model architecture that was previously designed based on a long short-term memory (LSTM)-based attention encoder decoder by applying transformer architectures. Secondly, we propose a speaker deduplication mechanism to reduce speaker identification errors in highly overlapped regions. Experimental results on the LibriSpeechMix dataset shows that the transformer-based architecture is especially good at counting the speakers and that the proposed model reduces the speaker-attributed word error rate by 47% over the LSTM-based baseline. Furthermore, for the LibriCSS dataset, which consists of real recordings of overlapped speech, the proposed model achieves concatenated minimum-permutation word error rates of 11.9% and 16.3% with and without target speaker profiles, respectively, both of which are the state-of-the-art results for LibriCSS with the monaural setting.

Comments:	Submitted to INTERSPEECH 2021
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
Cite as:	arXiv:2104.02128 [eess.AS]
	(or arXiv:2104.02128v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2104.02128

Submission history

From: Naoyuki Kanda [view email]
[v1] Mon, 5 Apr 2021 19:54:15 UTC (133 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:End-to-End Speaker-Attributed ASR with Transformer

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:End-to-End Speaker-Attributed ASR with Transformer

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators