Audio-Visual Model Distillation Using Acoustic Images

Pérez, Andrés F.; Sanguineti, Valentina; Morerio, Pietro; Murino, Vittorio

Computer Science > Computer Vision and Pattern Recognition

arXiv:1904.07933 (cs)

[Submitted on 16 Apr 2019 (v1), last revised 11 Feb 2020 (this version, v2)]

Title:Audio-Visual Model Distillation Using Acoustic Images

Authors:Andrés F. Pérez, Valentina Sanguineti, Pietro Morerio, Vittorio Murino

View PDF

Abstract:In this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality. Former models learn audio representations from raw signals or spectral data acquired by a single microphone, with remarkable results in classification and retrieval. However, such representations are not so robust towards variable environmental sound conditions. We tackle this drawback by exploiting a new multimodal labeled action recognition dataset acquired by a hybrid audio-visual sensor that provides RGB video, raw audio signals, and spatialized acoustic data, also known as acoustic images, where the visual and acoustic images are aligned in space and synchronized in time. Using this richer information, we train audio deep learning models in a teacher-student fashion. In particular, we distill knowledge into audio networks from both visual and acoustic image teachers. Our experiments suggest that the learned representations are more powerful and have better generalization capabilities than the features learned from models trained using just single-microphone audio data.

Comments:	Accepted at WACV 2020; supplementary material at page 11; code available at this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:1904.07933 [cs.CV]
	(or arXiv:1904.07933v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1904.07933

Submission history

From: Andrés F. Pérez [view email]
[v1] Tue, 16 Apr 2019 19:15:00 UTC (3,952 KB)
[v2] Tue, 11 Feb 2020 11:01:28 UTC (3,952 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Audio-Visual Model Distillation Using Acoustic Images

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Audio-Visual Model Distillation Using Acoustic Images

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators