Representation Learning on Visual-Symbolic Graphs for Video Understanding

Mavroudi, Effrosyni; Haro, Benjamín Béjar; Vidal, René

Computer Science > Computer Vision and Pattern Recognition

arXiv:1905.07385 (cs)

[Submitted on 17 May 2019 (v1), last revised 30 Sep 2020 (this version, v2)]

Title:Representation Learning on Visual-Symbolic Graphs for Video Understanding

Authors:Effrosyni Mavroudi, Benjamín Béjar Haro, René Vidal

View PDF

Abstract:Events in natural videos typically arise from spatio-temporal interactions between actors and objects and involve multiple co-occurring activities and object classes. To capture this rich visual and semantic context, we propose using two graphs: (1) an attributed spatio-temporal visual graph whose nodes correspond to actors and objects and whose edges encode different types of interactions, and (2) a symbolic graph that models semantic relationships. We further propose a graph neural network for refining the representations of actors, objects and their interactions on the resulting hybrid graph. Our model goes beyond current approaches that assume nodes and edges are of the same type, operate on graphs with fixed edge weights and do not use a symbolic graph. In particular, our framework: a) has specialized attention-based message functions for different node and edge types; b) uses visual edge features; c) integrates visual evidence with label relationships; and d) performs global reasoning in the semantic space. Experiments on challenging video understanding tasks, such as temporal action localization on the Charades dataset, show that the proposed method leads to state-of-the-art performance.

Comments:	ECCV 2020
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1905.07385 [cs.CV]
	(or arXiv:1905.07385v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1905.07385

Submission history

From: Effrosyni Mavroudi [view email]
[v1] Fri, 17 May 2019 17:33:48 UTC (1,974 KB)
[v2] Wed, 30 Sep 2020 15:55:51 UTC (1,987 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Representation Learning on Visual-Symbolic Graphs for Video Understanding

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Representation Learning on Visual-Symbolic Graphs for Video Understanding

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators