GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration

Sun, Yuchen; Zhao, Shanhui; Yu, Tao; Wen, Hao; Va, Samith; Xu, Mengwei; Li, Yuanchun; Zhang, Chongyang

Computer Science > Computer Vision and Pattern Recognition

arXiv:2503.17709 (cs)

[Submitted on 22 Mar 2025]

Title:GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration

Authors:Yuchen Sun, Shanhui Zhao, Tao Yu, Hao Wen, Samith Va, Mengwei Xu, Yuanchun Li, Chongyang Zhang

View PDF HTML (experimental)

Abstract:GUI agents hold significant potential to enhance the experience and efficiency of human-device interaction. However, current methods face challenges in generalizing across applications (apps) and tasks, primarily due to two fundamental limitations in existing datasets. First, these datasets overlook developer-induced structural variations among apps, limiting the transferability of knowledge across diverse software environments. Second, many of them focus solely on navigation tasks, which restricts their capacity to represent comprehensive software architectures and complex user interactions. To address these challenges, we introduce GUI-Xplore, a dataset meticulously designed to enhance cross-application and cross-task generalization via an exploration-and-reasoning framework. GUI-Xplore integrates pre-recorded exploration videos providing contextual insights, alongside five hierarchically structured downstream tasks designed to comprehensively evaluate GUI agent capabilities. To fully exploit GUI-Xplore's unique features, we propose Xplore-Agent, a GUI agent framework that combines Action-aware GUI Modeling with Graph-Guided Environment Reasoning. Further experiments indicate that Xplore-Agent achieves a 10% improvement over existing methods in unfamiliar environments, yet there remains significant potential for further enhancement towards truly generalizable GUI agents.

Comments:	CVPR 2025
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2503.17709 [cs.CV]
	(or arXiv:2503.17709v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2503.17709

Submission history

From: Yuchen Sun [view email]
[v1] Sat, 22 Mar 2025 09:30:37 UTC (22,047 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators