Finding Fallen Objects Via Asynchronous Audio-Visual Integration

Gan, Chuang; Gu, Yi; Zhou, Siyuan; Schwartz, Jeremy; Alter, Seth; Traer, James; Gutfreund, Dan; Tenenbaum, Joshua B.; McDermott, Josh; Torralba, Antonio

Computer Science > Computer Vision and Pattern Recognition

arXiv:2207.03483 (cs)

[Submitted on 7 Jul 2022]

Title:Finding Fallen Objects Via Asynchronous Audio-Visual Integration

Authors:Chuang Gan, Yi Gu, Siyuan Zhou, Jeremy Schwartz, Seth Alter, James Traer, Dan Gutfreund, Joshua B. Tenenbaum, Josh McDermott, Antonio Torralba

View PDF

Abstract:The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and then must find it. In this paper, we introduce a setting in which to study multi-modal object localization in 3D virtual environments. An object is dropped somewhere in a room. An embodied robot agent, equipped with a camera and microphone, must determine what object has been dropped -- and where -- by combining audio and visual signals with knowledge of the underlying physics. To study this problem, we have generated a large-scale dataset -- the Fallen Objects dataset -- that includes 8000 instances of 30 physical object categories in 64 rooms. The dataset uses the ThreeDWorld platform which can simulate physics-based impact sounds and complex physical interactions between objects in a photorealistic setting. As a first step toward addressing this challenge, we develop a set of embodied agent baselines, based on imitation learning, reinforcement learning, and modular planning, and perform an in-depth analysis of the challenge of this new task.

Comments:	CVPR 2022. Project page: this http URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2207.03483 [cs.CV]
	(or arXiv:2207.03483v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2207.03483

Submission history

From: Chuang Gan [view email]
[v1] Thu, 7 Jul 2022 17:59:59 UTC (30,410 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Finding Fallen Objects Via Asynchronous Audio-Visual Integration

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Finding Fallen Objects Via Asynchronous Audio-Visual Integration

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators