Binary Image Selection (BISON): Interpretable Evaluation of Visual Grounding

Hu, Hexiang; Misra, Ishan; van der Maaten, Laurens

Computer Science > Computer Vision and Pattern Recognition

arXiv:1901.06595v1 (cs)

[Submitted on 19 Jan 2019 (this version), latest version 5 Apr 2019 (v2)]

Title:Binary Image Selection (BISON): Interpretable Evaluation of Visual Grounding

Authors:Hexiang Hu, Ishan Misra, Laurens van der Maaten

View PDF

Abstract:Providing systems the ability to relate linguistic and visual content is one of the hallmarks of computer vision. Tasks such as image captioning and retrieval were designed to test this ability, but come with complex evaluation measures that gauge various other abilities and biases simultaneously. This paper presents an alternative evaluation task for visual-grounding systems: given a caption the system is asked to select the image that best matches the caption from a pair of semantically similar images. The system's accuracy on this Binary Image SelectiON (BISON) task is not only interpretable, but also measures the ability to relate fine-grained text content in the caption to visual content in the images. We gathered a BISON dataset that complements the COCO Captions dataset and used this dataset in auxiliary evaluations of captioning and caption-based retrieval systems. While captioning measures suggest visual grounding systems outperform humans, BISON shows that these systems are still far away from human performance.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:1901.06595 [cs.CV]
	(or arXiv:1901.06595v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1901.06595

Submission history

From: Hexiang Hu [view email]
[v1] Sat, 19 Jan 2019 22:12:01 UTC (8,239 KB)
[v2] Fri, 5 Apr 2019 16:34:48 UTC (8,152 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CV

< prev | next >

new | recent | 2019-01

Change to browse by:

cs
cs.AI
cs.CL

References & Citations

DBLP - CS Bibliography

listing | bibtex

Hexiang Hu
Ishan Misra
Laurens van der Maaten

export BibTeX citation

Computer Science > Computer Vision and Pattern Recognition

Title:Binary Image Selection (BISON): Interpretable Evaluation of Visual Grounding

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Binary Image Selection (BISON): Interpretable Evaluation of Visual Grounding

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators