Dual Normalization Multitasking for Audio-Visual Sounding Object Localization

Nishikawa, Tokuhiro; Shimada, Daiki; Yokono, Jerry Jun

Computer Science > Computer Vision and Pattern Recognition

arXiv:2106.00180 (cs)

[Submitted on 1 Jun 2021]

Title:Dual Normalization Multitasking for Audio-Visual Sounding Object Localization

Authors:Tokuhiro Nishikawa, Daiki Shimada, Jerry Jun Yokono

View PDF

Abstract:Although several research works have been reported on audio-visual sound source localization in unconstrained videos, no datasets and metrics have been proposed in the literature to quantitatively evaluate its performance. Defining the ground truth for sound source localization is difficult, because the location where the sound is produced is not limited to the range of the source object, but the vibrations propagate and spread through the surrounding objects. Therefore we propose a new concept, Sounding Object, to reduce the ambiguity of the visual location of sound, making it possible to annotate the location of the wide range of sound sources. With newly proposed metrics for quantitative evaluation, we formulate the problem of Audio-Visual Sounding Object Localization (AVSOL). We also created the evaluation dataset (AVSOL-E dataset) by manually annotating the test set of well-known Audio-Visual Event (AVE) dataset. To tackle this new AVSOL problem, we propose a novel multitask training strategy and architecture called Dual Normalization Multitasking (DNM), which aggregates the Audio-Visual Correspondence (AVC) task and the classification task for video events into a single audio-visual similarity map. By efficiently utilize both supervisions by DNM, our proposed architecture significantly outperforms the baseline methods.

Comments:	10 pages, 6 figures
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2106.00180 [cs.CV]
	(or arXiv:2106.00180v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2106.00180

Submission history

From: Tokuhiro Nishikawa [view email]
[v1] Tue, 1 Jun 2021 02:02:52 UTC (11,534 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Dual Normalization Multitasking for Audio-Visual Sounding Object Localization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Dual Normalization Multitasking for Audio-Visual Sounding Object Localization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators