A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection

Deng, Qingkun; Luz, Saturnino; Garcia, Sofia de la Fuente

Computer Science > Sound

arXiv:2406.03138v1 (cs)

[Submitted on 5 Jun 2024 (this version), latest version 19 May 2025 (v3)]

Title:A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection

Authors:Qingkun Deng, Saturnino Luz, Sofia de la Fuente Garcia

View PDF HTML (experimental)

Abstract:Speech-based depression detection tools could help early screening of depression. Here, we address two issues that may hinder the clinical practicality of such tools: segment-level labelling noise and a lack of model interpretability. We propose a speech-level Audio Spectrogram Transformer to avoid segment-level labelling. We observe that the proposed model significantly outperforms a segment-level model, providing evidence for the presence of segment-level labelling noise in audio modality and the advantage of longer-duration speech analysis for depression detection. We introduce a frame-based attention interpretation method to extract acoustic features from prediction-relevant waveform signals for interpretation by clinicians. Through interpretation, we observe that the proposed model identifies reduced loudness and F0 as relevant signals of depression, which aligns with the speech characteristics of depressed patients documented in clinical studies.

Comments:	5 pages, 3 figures. arXiv admin note: substantial text overlap with arXiv:2309.13476
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2406.03138 [cs.SD]
	(or arXiv:2406.03138v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2406.03138

Submission history

From: Qingkun Deng [view email]
[v1] Wed, 5 Jun 2024 10:47:00 UTC (1,540 KB)
[v2] Fri, 7 Jun 2024 09:27:16 UTC (1,191 KB)
[v3] Mon, 19 May 2025 11:51:34 UTC (2,180 KB)

Computer Science > Sound

Title:A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators