x-vectors meet emotions: A study on dependencies between emotion and speaker recognition

Pappagari, Raghavendra; Wang, Tianzi; Villalba, Jesus; Chen, Nanxin; Dehak, Najim

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2002.05039 (eess)

[Submitted on 12 Feb 2020]

Title:x-vectors meet emotions: A study on dependencies between emotion and speaker recognition

Authors:Raghavendra Pappagari, Tianzi Wang, Jesus Villalba, Nanxin Chen, Najim Dehak

View PDF

Abstract:In this work, we explore the dependencies between speaker recognition and emotion recognition. We first show that knowledge learned for speaker recognition can be reused for emotion recognition through transfer learning. Then, we show the effect of emotion on speaker recognition. For emotion recognition, we show that using a simple linear model is enough to obtain good performance on the features extracted from pre-trained models such as the x-vector model. Then, we improve emotion recognition performance by fine-tuning for emotion classification. We evaluated our experiments on three different types of datasets: IEMOCAP, MSP-Podcast, and Crema-D. By fine-tuning, we obtained 30.40%, 7.99%, and 8.61% absolute improvement on IEMOCAP, MSP-Podcast, and Crema-D respectively over baseline model with no pre-training. Finally, we present results on the effect of emotion on speaker verification. We observed that speaker verification performance is prone to changes in test speaker emotions. We found that trials with angry utterances performed worst in all three datasets. We hope our analysis will initiate a new line of research in the speaker recognition community.

Comments:	45th International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2020
Subjects:	Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
Cite as:	arXiv:2002.05039 [eess.AS]
	(or arXiv:2002.05039v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2002.05039

Submission history

From: Raghavendra Reddy Pappagari [view email]
[v1] Wed, 12 Feb 2020 15:13:07 UTC (45 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:x-vectors meet emotions: A study on dependencies between emotion and speaker recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:x-vectors meet emotions: A study on dependencies between emotion and speaker recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators