MObI: Multimodal Object Inpainting Using Diffusion Models

Buburuzan, Alexandru; Sharma, Anuj; Redford, John; Dokania, Puneet K.; Mueller, Romain

Computer Science > Computer Vision and Pattern Recognition

arXiv:2501.03173 (cs)

[Submitted on 6 Jan 2025 (v1), last revised 22 Apr 2025 (this version, v2)]

Title:MObI: Multimodal Object Inpainting Using Diffusion Models

Authors:Alexandru Buburuzan, Anuj Sharma, John Redford, Puneet K. Dokania, Romain Mueller

View PDF HTML (experimental)

Abstract:Safety-critical applications, such as autonomous driving, require extensive multimodal data for rigorous testing. Methods based on synthetic data are gaining prominence due to the cost and complexity of gathering real-world data but require a high degree of realism and controllability in order to be useful. This paper introduces MObI, a novel framework for Multimodal Object Inpainting that leverages a diffusion model to create realistic and controllable object inpaintings across perceptual modalities, demonstrated for both camera and lidar simultaneously. Using a single reference RGB image, MObI enables objects to be seamlessly inserted into existing multimodal scenes at a 3D location specified by a bounding box, while maintaining semantic consistency and multimodal coherence. Unlike traditional inpainting methods that rely solely on edit masks, our 3D bounding box conditioning gives objects accurate spatial positioning and realistic scaling. As a result, our approach can be used to insert novel objects flexibly into multimodal scenes, providing significant advantages for testing perception models.

Comments:	8 pages; Project page at this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2501.03173 [cs.CV]
	(or arXiv:2501.03173v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2501.03173

Submission history

From: Romain Mueller [view email]
[v1] Mon, 6 Jan 2025 17:43:26 UTC (14,855 KB)
[v2] Tue, 22 Apr 2025 11:09:48 UTC (16,695 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:MObI: Multimodal Object Inpainting Using Diffusion Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:MObI: Multimodal Object Inpainting Using Diffusion Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators