Vision-Language Pre-Training with Triple Contrastive Learning

Yang, Jinyu; Duan, Jiali; Tran, Son; Xu, Yi; Chanda, Sampath; Chen, Liqun; Zeng, Belinda; Chilimbi, Trishul; Huang, Junzhou

Computer Science > Computer Vision and Pattern Recognition

arXiv:2202.10401 (cs)

[Submitted on 21 Feb 2022 (v1), last revised 28 Mar 2022 (this version, v4)]

Title:Vision-Language Pre-Training with Triple Contrastive Learning

Authors:Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, Junzhou Huang

View PDF

Abstract:Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information (MI) between an image and its matched text. However, simply performing cross-modal alignment (CMA) ignores data potential within each modality, which may result in degraded representations. For instance, although CMA-based models are able to map image-text pairs close together in the embedding space, they fail to ensure that similar inputs from the same modality stay close by. This problem can get even worse when the pre-training data is noisy. In this paper, we propose triple contrastive learning (TCL) for vision-language pre-training by leveraging both cross-modal and intra-modal self-supervision. Besides CMA, TCL introduces an intra-modal contrastive objective to provide complementary benefits in representation learning. To take advantage of localized and structural information from image and text input, TCL further maximizes the average MI between local regions of image/text and their global summary. To the best of our knowledge, ours is the first work that takes into account local structure information for multi-modality representation learning. Experimental evaluations show that our approach is competitive and achieves the new state of the art on various common down-stream vision-language tasks such as image-text retrieval and visual question answering.

Comments:	CVPR 2022; code: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2202.10401 [cs.CV]
	(or arXiv:2202.10401v4 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2202.10401

Submission history

From: Jinyu Yang [view email]
[v1] Mon, 21 Feb 2022 17:54:57 UTC (1,932 KB)
[v2] Wed, 2 Mar 2022 16:20:34 UTC (1,932 KB)
[v3] Thu, 3 Mar 2022 05:15:25 UTC (1,932 KB)
[v4] Mon, 28 Mar 2022 14:44:39 UTC (1,016 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Vision-Language Pre-Training with Triple Contrastive Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Vision-Language Pre-Training with Triple Contrastive Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators