An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech

Deng, Qingkun; Luz, Saturnino; Garcia, Sofia de la Fuente

Computer Science > Sound

arXiv:2406.03138 (cs)

[Submitted on 5 Jun 2024 (v1), last revised 19 May 2025 (this version, v3)]

Title:An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech

Authors:Qingkun Deng, Saturnino Luz, Sofia de la Fuente Garcia

View PDF HTML (experimental)

Abstract:Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram Transformer (AST) to detect depression using long-duration speech instead of short segments, along with a novel interpretation method that reveals prediction-relevant acoustic features for clinician interpretation. Our experiments show the proposed model outperforms a segment-level AST, highlighting the impact of segment-level labelling noise and the advantage of leveraging longer speech duration for more reliable depression detection. Through interpretation, we observe our model identifies reduced loudness and F0 as relevant depression signals, aligning with documented clinical findings. This interpretability supports a responsible AI approach for speech-based depression detection, rendering such tools more clinically applicable.

Comments:	5 pages, 3 figures. arXiv admin note: substantial text overlap with arXiv:2309.13476
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2406.03138 [cs.SD]
	(or arXiv:2406.03138v3 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2406.03138

Submission history

From: Qingkun Deng [view email]
[v1] Wed, 5 Jun 2024 10:47:00 UTC (1,540 KB)
[v2] Fri, 7 Jun 2024 09:27:16 UTC (1,191 KB)
[v3] Mon, 19 May 2025 11:51:34 UTC (2,180 KB)

Computer Science > Sound

Title:An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators