Joint framework with deep feature distillation and adaptive focal loss for weakly supervised audio tagging and acoustic event detection

Liang, Yunhao; Long, Yanhua; Li, Yijie; Liang, Jiaen; Wang, Yuping

doi:10.1016/j.dsp.2022.103446

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2103.12388 (eess)

[Submitted on 23 Mar 2021 (v1), last revised 12 Feb 2022 (this version, v2)]

Title:Joint framework with deep feature distillation and adaptive focal loss for weakly supervised audio tagging and acoustic event detection

Authors:Yunhao Liang, Yanhua Long, Yijie Li, Jiaen Liang, Yuping Wang

View PDF

Abstract:A good joint training framework is very helpful to improve the performances of weakly supervised audio tagging (AT) and acoustic event detection (AED) simultaneously. In this study, we propose three methods to improve the best teacher-student framework in the IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) 2019 Task 4 for both audio tagging and acoustic events detection tasks. A frame-level target-events based deep feature distillation is first proposed, which aims to leverage the potential of limited strong-labeled data in weakly supervised framework to learn better intermediate feature maps. Then, we propose an adaptive focal loss and two-stage training strategy to enable an effective and more accurate model training, where the contribution of hard and easy acoustic events to the total cost function can be automatically adjusted. Furthermore, an event-specific post processing is designed to improve the prediction of target event time-stamps. Our experiments are performed on the public DCASE 2019 Task 4 dataset, results show that our approach achieves competitive performances in both AT (81.2\% F1-score) and AED (49.8\% F1-score) tasks.

Comments:	Updated, please refer to "this https URL
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2103.12388 [eess.AS]
	(or arXiv:2103.12388v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2103.12388
Related DOI:	https://doi.org/10.1016/j.dsp.2022.103446

Submission history

From: Yunhao Liang [view email]
[v1] Tue, 23 Mar 2021 08:44:07 UTC (478 KB)
[v2] Sat, 12 Feb 2022 09:00:34 UTC (1,078 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Joint framework with deep feature distillation and adaptive focal loss for weakly supervised audio tagging and acoustic event detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Joint framework with deep feature distillation and adaptive focal loss for weakly supervised audio tagging and acoustic event detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators