Curriculum Learning for Data-Efficient Vision-Language Alignment

Srinivasan, Tejas; Ren, Xiang; Thomason, Jesse

Computer Science > Computer Vision and Pattern Recognition

arXiv:2207.14525 (cs)

[Submitted on 29 Jul 2022]

Title:Curriculum Learning for Data-Efficient Vision-Language Alignment

Authors:Tejas Srinivasan, Xiang Ren, Jesse Thomason

View PDF

Abstract:Aligning image and text encoders from scratch using contrastive learning requires large amounts of paired image-text data. We alleviate this need by aligning individually pre-trained language and vision representation models using a much smaller amount of paired data, augmented with a curriculum learning algorithm to learn fine-grained vision-language alignments. TOnICS (Training with Ontology-Informed Contrastive Sampling) initially samples minibatches whose image-text pairs contain a wide variety of objects to learn object-level alignment, and progressively samples minibatches where all image-text pairs contain the same object to learn finer-grained contextual alignment. Aligning pre-trained BERT and VinVL models to each other using TOnICS outperforms CLIP on downstream zero-shot image retrieval while using less than 1% as much training data.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2207.14525 [cs.CV]
	(or arXiv:2207.14525v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2207.14525

Submission history

From: Tejas Srinivasan [view email]
[v1] Fri, 29 Jul 2022 07:45:56 UTC (8,571 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Curriculum Learning for Data-Efficient Vision-Language Alignment

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Curriculum Learning for Data-Efficient Vision-Language Alignment

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators