Speech Emotion Recognition using Multi-task learning and a multimodal dynamic fusion network

Ghosh, Sreyan; Ramaneswaran, S; Srivastava, Harshvardhan; Umesh, S.

Computer Science > Computation and Language

arXiv:2203.16794v4 (cs)

[Submitted on 31 Mar 2022 (v1), revised 31 Oct 2022 (this version, v4), latest version 3 Jun 2023 (v5)]

Title:Speech Emotion Recognition using Multi-task learning and a multimodal dynamic fusion network

Authors:Sreyan Ghosh, S Ramaneswaran, Harshvardhan Srivastava, S. Umesh

View PDF

Abstract:Emotion Recognition (ER) aims to classify human utterances into different emotion categories. Based on early-fusion and self-attention-based multimodal interaction between text and acoustic modalities, in this paper, we propose MMER, a multimodal multitask learning approach for ER from individual utterances in isolation. Our proposed MMER leverages a multimodal dynamic fusion network that adds minimal parameters over an existing speech encoder to leverage the semantic and syntactic properties hidden in text. Experiments on the IEMOCAP benchmark show that our proposed model achieves state-of-the-art performance. In addition, strong baselines and ablation studies prove the effectiveness of our proposed approach. We make our code publicly available on GitHub.

Comments:	6 + 2 pages
Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2203.16794 [cs.CL]
	(or arXiv:2203.16794v4 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2203.16794

Submission history

From: Sreyan Ghosh [view email]
[v1] Thu, 31 Mar 2022 04:51:32 UTC (312 KB)
[v2] Fri, 1 Apr 2022 04:39:53 UTC (312 KB)
[v3] Thu, 18 Aug 2022 15:12:39 UTC (312 KB)
[v4] Mon, 31 Oct 2022 21:51:49 UTC (1,446 KB)
[v5] Sat, 3 Jun 2023 21:55:28 UTC (711 KB)

Computer Science > Computation and Language

Title:Speech Emotion Recognition using Multi-task learning and a multimodal dynamic fusion network

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Speech Emotion Recognition using Multi-task learning and a multimodal dynamic fusion network

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators