C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

Wang, Siheng; Li, Zhengdao; Li, Yanshu; Xiao, Canran; Zhan, Haibo; Yao, Zhengtao; Zhang, Xuzhi; Kang, Jiale; Li, Linshan; Liu, Weiming; Dong, Zhikang; Shen, Jifeng; Dong, Junhao; Sun, Qiang; Koniusz, Piotr

Computer Science > Computer Vision and Pattern Recognition

arXiv:2509.23316 (cs)

[Submitted on 27 Sep 2025]

Title:C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

Authors:Siheng Wang, Zhengdao Li, Yanshu Li, Canran Xiao, Haibo Zhan, Zhengtao Yao, Xuzhi Zhang, Jiale Kang, Linshan Li, Weiming Liu, Zhikang Dong, Jifeng Shen, Junhao Dong, Qiang Sun, Piotr Koniusz

View PDF HTML (experimental)

Abstract:Object detection has advanced significantly in the closed-set setting, but real-world deployment remains limited by two challenges: poor generalization to unseen categories and insufficient robustness under adverse conditions. Prior research has explored these issues separately: visible-infrared detection improves robustness but lacks generalization, while open-world detection leverages vision-language alignment strategy for category diversity but struggles under extreme environments. This trade-off leaves robustness and diversity difficult to achieve simultaneously. To mitigate these issues, we propose \textbf{C3-OWD}, a curriculum cross-modal contrastive learning framework that unifies both strengths. Stage~1 enhances robustness by pretraining with RGBT data, while Stage~2 improves generalization via vision-language alignment. To prevent catastrophic forgetting between two stages, we introduce an Exponential Moving Average (EMA) mechanism that theoretically guarantees preservation of pre-stage performance with bounded parameter lag and function consistency. Experiments on FLIR, OV-COCO, and OV-LVIS demonstrate the effectiveness of our approach: C3-OWD achieves $80.1$ AP$^{50}$ on FLIR, $48.6$ AP$^{50}_{\text{Novel}}$ on OV-COCO, and $35.7$ mAP$_r$ on OV-LVIS, establishing competitive performance across both robustness and diversity evaluations. Code available at: this https URL.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2509.23316 [cs.CV]
	(or arXiv:2509.23316v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2509.23316

Submission history

From: Siheng Wang [view email]
[v1] Sat, 27 Sep 2025 14:04:15 UTC (1,715 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators