Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Zhao, Zongchuang; Fu, Haoyu; Liang, Dingkang; Zhou, Xin; Zhang, Dingyuan; Xie, Hongwei; Wang, Bing; Bai, Xiang

Computer Science > Computer Vision and Pattern Recognition

arXiv:2505.08725 (cs)

[Submitted on 13 May 2025]

Title:Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Authors:Zongchuang Zhao, Haoyu Fu, Dingkang Liang, Xin Zhou, Dingyuan Zhang, Hongwei Xie, Bing Wang, Xiang Bai

View PDF HTML (experimental)

Abstract:The Large Visual-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing research typically focuses on front-view perspectives and partial objects within scenes, struggling to achieve comprehensive scene understanding. Meanwhile, existing LVLMs suffer from the lack of mapping relationship between 2D and 3D and insufficient integration of 3D object localization and instruction understanding. To tackle these limitations, we first introduce NuInteract, a large-scale dataset with over 1.5M multi-view image language pairs spanning dense scene captions and diverse interactive tasks. Furthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries. The spatial processor, designed as a plug-and-play component, can be initialized with pre-trained 3D detectors to improve 3D perception. Our experiments show that DriveMonkey outperforms general LVLMs, especially achieving a 9.86% notable improvement on the 3D visual grounding task. The dataset and code will be released at this https URL.

Comments:	The dataset and code will be released at this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2505.08725 [cs.CV]
	(or arXiv:2505.08725v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2505.08725

Submission history

From: Zongchuang Zhao [view email]
[v1] Tue, 13 May 2025 16:36:51 UTC (837 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators