SimGen: Simulator-conditioned Driving Scene Generation

Zhou, Yunsong; Simon, Michael; Peng, Zhenghao; Mo, Sicheng; Zhu, Hongzi; Guo, Minyi; Zhou, Bolei

Computer Science > Computer Vision and Pattern Recognition

arXiv:2406.09386 (cs)

[Submitted on 13 Jun 2024 (v1), last revised 7 Dec 2024 (this version, v3)]

Title:SimGen: Simulator-conditioned Driving Scene Generation

Authors:Yunsong Zhou, Michael Simon, Zhenghao Peng, Sicheng Mo, Hongzi Zhu, Minyi Guo, Bolei Zhou

View PDF HTML (experimental)

Abstract:Controllable synthetic data generation can substantially lower the annotation cost of training data. Prior works use diffusion models to generate driving images conditioned on the 3D object layout. However, those models are trained on small-scale datasets like nuScenes, which lack appearance and layout diversity. Moreover, overfitting often happens, where the trained models can only generate images based on the layout data from the validation set of the same dataset. In this work, we introduce a simulator-conditioned scene generation framework called SimGen that can learn to generate diverse driving scenes by mixing data from the simulator and the real world. It uses a novel cascade diffusion pipeline to address challenging sim-to-real gaps and multi-condition conflicts. A driving video dataset DIVA is collected to enhance the generative diversity of SimGen, which contains over 147.5 hours of real-world driving videos from 73 locations worldwide and simulated driving data from the MetaDrive simulator. SimGen achieves superior generation quality and diversity while preserving controllability based on the text prompt and the layout pulled from a simulator. We further demonstrate the improvements brought by SimGen for synthetic data augmentation on the BEV detection and segmentation task and showcase its capability in safety-critical data generation.

Comments:	arXiv admin note: text overlap with arXiv:2403.09630 by other authors
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2406.09386 [cs.CV]
	(or arXiv:2406.09386v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2406.09386

Submission history

From: Yunsong Zhou [view email]
[v1] Thu, 13 Jun 2024 17:58:32 UTC (33,018 KB)
[v2] Mon, 28 Oct 2024 07:19:45 UTC (37,895 KB)
[v3] Sat, 7 Dec 2024 10:25:34 UTC (37,298 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:SimGen: Simulator-conditioned Driving Scene Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:SimGen: Simulator-conditioned Driving Scene Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators