ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Liu, Ruiping; Zhang, Jiaming; Schön, Angela; Müller, Karin; Zheng, Junwei; Yang, Kailun; Guo, Anhong; Gerling, Kathrin; Stiefelhagen, Rainer

Computer Science > Human-Computer Interaction

arXiv:2412.03118 (cs)

[Submitted on 4 Dec 2024 (v1), last revised 30 Apr 2025 (this version, v2)]

Title:ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Authors:Ruiping Liu, Jiaming Zhang, Angela Schön, Karin Müller, Junwei Zheng, Kailun Yang, Anhong Guo, Kathrin Gerling, Rainer Stiefelhagen

View PDF HTML (experimental)

Abstract:Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing description- and detection-based assistive technologies do not sufficiently support the multifaceted nature of interactive object search tasks. We present ObjectFinder, an open-vocabulary wearable assistive system for interactive object search by blind people. ObjectFinder allows users to query target objects using flexible wording. Once the target object is detected, it provides egocentric localization information in real-time, including distance and direction. Users can then initiate different branches to gather detailed information based on their intent towards the target object, such as navigating to it or perceiving its surroundings. ObjectFinder is powered by a seamless combination of open-vocabulary models, namely an open-vocabulary object detector and a multimodal large language model. The ObjectFinder design concept and its development were carried out in collaboration with a blind co-designer. To evaluate ObjectFinder, we conducted an exploratory user study with eight blind participants. We compared ObjectFinder to BeMyAI and Google Lookout, popular description- and detection-based assistive applications. Our findings indicate that most participants felt more independent with ObjectFinder and preferred it for object search, as it enhanced scene context gathering and navigation, and allowed for active target identification. Finally, we discuss the implications for future assistive systems to support interactive object search.

Subjects:	Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2412.03118 [cs.HC]
	(or arXiv:2412.03118v2 [cs.HC] for this version)
	https://doi.org/10.48550/arXiv.2412.03118

Submission history

From: Jiaming Zhang [view email]
[v1] Wed, 4 Dec 2024 08:38:45 UTC (37,411 KB)
[v2] Wed, 30 Apr 2025 17:42:40 UTC (47,720 KB)

Computer Science > Human-Computer Interaction

Title:ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Human-Computer Interaction

Title:ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators