Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition

Lin, Kun-Yu; Ding, Henghui; Zhou, Jiaming; Tang, Yu-Ming; Peng, Yi-Xing; Zhao, Zhilin; Loy, Chen Change; Zheng, Wei-Shi

Computer Science > Computer Vision and Pattern Recognition

arXiv:2403.01560 (cs)

[Submitted on 3 Mar 2024 (v1), last revised 24 May 2024 (this version, v2)]

Title:Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition

Authors:Kun-Yu Lin, Henghui Ding, Jiaming Zhou, Yu-Ming Tang, Yi-Xing Peng, Zhilin Zhao, Chen Change Loy, Wei-Shi Zheng

View PDF HTML (experimental)

Abstract:Building upon the impressive success of CLIP (Contrastive Language-Image Pretraining), recent pioneer works have proposed to adapt the powerful CLIP to video data, leading to efficient and effective video learners for open-vocabulary action recognition. Inspired by that humans perform actions in diverse environments, our work delves into an intriguing question: Can CLIP-based video learners effectively generalize to video domains they have not encountered during training? To answer this, we establish a CROSS-domain Open-Vocabulary Action recognition benchmark named XOV-Action, and conduct a comprehensive evaluation of five state-of-the-art CLIP-based video learners under various types of domain gaps. The evaluation demonstrates that previous methods exhibit limited action recognition performance in unseen video domains, revealing potential challenges of the cross-domain open-vocabulary action recognition task. In this paper, we focus on one critical challenge of the task, namely scene bias, and accordingly contribute a novel scene-aware video-text alignment method. Our key idea is to distinguish video representations apart from scene-encoded text representations, aiming to learn scene-agnostic video representations for recognizing actions across domains. Extensive experiments demonstrate the effectiveness of our method. The benchmark and code will be available at this https URL.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2403.01560 [cs.CV]
	(or arXiv:2403.01560v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2403.01560

Submission history

From: Kun-Yu Lin [view email]
[v1] Sun, 3 Mar 2024 16:48:16 UTC (370 KB)
[v2] Fri, 24 May 2024 14:47:03 UTC (617 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators