Weakly-Supervised Temporal Article Grounding

Chen, Long; Niu, Yulei; Chen, Brian; Lin, Xudong; Han, Guangxing; Thomas, Christopher; Ayyubi, Hammad; Ji, Heng; Chang, Shih-Fu

Computer Science > Computer Vision and Pattern Recognition

arXiv:2210.12444 (cs)

[Submitted on 22 Oct 2022 (v1), last revised 24 Feb 2023 (this version, v2)]

Title:Weakly-Supervised Temporal Article Grounding

Authors:Long Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, Shih-Fu Chang

View PDF

Abstract:Given a long untrimmed video and natural language queries, video grounding (VG) aims to temporally localize the semantically-aligned video segments. Almost all existing VG work holds two simple but unrealistic assumptions: 1) All query sentences can be grounded in the corresponding video. 2) All query sentences for the same video are always at the same semantic scale. Unfortunately, both assumptions make today's VG models fail to work in practice. For example, in real-world multimodal assets (eg, news articles), most of the sentences in the article can not be grounded in their affiliated videos, and they typically have rich hierarchical relations (ie, at different semantic scales). To this end, we propose a new challenging grounding task: Weakly-Supervised temporal Article Grounding (WSAG). Specifically, given an article and a relevant video, WSAG aims to localize all ``groundable'' sentences to the video, and these sentences are possibly at different semantic scales. Accordingly, we collect the first WSAG dataset to facilitate this task: YouwikiHow, which borrows the inherent multi-scale descriptions in wikiHow articles and plentiful YouTube videos. In addition, we propose a simple but effective method DualMIL for WSAG, which consists of a two-level MIL loss and a single-/cross- sentence constraint loss. These training objectives are carefully designed for these relaxed assumptions. Extensive ablations have verified the effectiveness of DualMIL.

Comments:	EMNLP 2022, this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
Cite as:	arXiv:2210.12444 [cs.CV]
	(or arXiv:2210.12444v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2210.12444

Submission history

From: Long Chen [view email]
[v1] Sat, 22 Oct 2022 13:23:02 UTC (9,061 KB)
[v2] Fri, 24 Feb 2023 02:53:39 UTC (7,806 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Weakly-Supervised Temporal Article Grounding

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Weakly-Supervised Temporal Article Grounding

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators