Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

Han, Ting; Zarrieß, Sina

Computer Science > Robotics

arXiv:2101.12338 (cs)

[Submitted on 14 Jan 2021]

Title:Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

Authors:Ting Han, Sina Zarrieß

View PDF

Abstract:Socially competent robots should be equipped with the ability to perceive the world that surrounds them and communicate about it in a human-like manner. Representative skills that exhibit such ability include generating image descriptions and visually grounded referring expressions. In the NLG community, these generation tasks are largely investigated in non-interactive and language-only settings. However, in face-to-face interaction, humans often deploy multiple modalities to communicate, forming seamless integration of natural language, hand gestures and other modalities like sketches. To enable robots to describe what they perceive with speech and sketches/gestures, we propose to model the task of generating natural language together with free-hand sketches/hand gestures to describe visual scenes and real life objects, namely, visually-grounded multimodal description generation. In this paper, we discuss the challenges and evaluation metrics of the task, and how the task can benefit from progress recently made in the natural language processing and computer vision realms, where related topics such as visually grounded NLG, distributional semantics, and photo-based sketch generation have been extensively studied.

Comments:	The 2nd Workshop on NLG for HRI colocated with The 13th International Conference on Natural Language Generation
Subjects:	Robotics (cs.RO); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2101.12338 [cs.RO]
	(or arXiv:2101.12338v1 [cs.RO] for this version)
	https://doi.org/10.48550/arXiv.2101.12338

Submission history

From: Ting Han [view email]
[v1] Thu, 14 Jan 2021 23:40:23 UTC (180 KB)

Computer Science > Robotics

Title:Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Robotics

Title:Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators