Referring Human Pose and Mask Estimation in the Wild

Miao, Bo; Feng, Mingtao; Wu, Zijie; Bennamoun, Mohammed; Gao, Yongsheng; Mian, Ajmal

Computer Science > Computer Vision and Pattern Recognition

arXiv:2410.20508 (cs)

[Submitted on 27 Oct 2024]

Title:Referring Human Pose and Mask Estimation in the Wild

Authors:Bo Miao, Mingtao Feng, Zijie Wu, Mohammed Bennamoun, Yongsheng Gao, Ajmal Mian

View PDF HTML (experimental)

Abstract:We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to previous works, R-HPM (i) ensures high-quality, identity-aware results corresponding to the referred person, and (ii) simultaneously predicts human pose and mask for a comprehensive representation. To achieve this, we introduce a large-scale dataset named RefHuman, which substantially extends the MS COCO dataset with additional text and positional prompt annotations. RefHuman includes over 50,000 annotated instances in the wild, each equipped with keypoint, mask, and prompt annotations. To enable prompt-conditioned estimation, we propose the first end-to-end promptable approach named UniPHD for R-HPM. UniPHD extracts multimodal representations and employs a proposed pose-centric hierarchical decoder to process (text or positional) instance queries and keypoint queries, producing results specific to the referred person. Extensive experiments demonstrate that UniPHD produces quality results based on user-friendly prompts and achieves top-tier performance on RefHuman val and MS COCO val2017. Data and Code: this https URL

Comments:	Accepted by NeurIPS 2024. this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
Cite as:	arXiv:2410.20508 [cs.CV]
	(or arXiv:2410.20508v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2410.20508

Submission history

From: Bo Miao [view email]
[v1] Sun, 27 Oct 2024 16:44:15 UTC (2,904 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Referring Human Pose and Mask Estimation in the Wild

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Referring Human Pose and Mask Estimation in the Wild

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators