Adversarial Nibbler: A Data-Centric Challenge for Improving the Safety of Text-to-Image Models

Parrish, Alicia; Kirk, Hannah Rose; Quaye, Jessica; Rastogi, Charvi; Bartolo, Max; Inel, Oana; Ciro, Juan; Mosquera, Rafael; Howard, Addison; Cukierski, Will; Sculley, D.; Reddi, Vijay Janapa; Aroyo, Lora

Computer Science > Machine Learning

arXiv:2305.14384 (cs)

[Submitted on 22 May 2023]

Title:Adversarial Nibbler: A Data-Centric Challenge for Improving the Safety of Text-to-Image Models

Authors:Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Max Bartolo, Oana Inel, Juan Ciro, Rafael Mosquera, Addison Howard, Will Cukierski, D. Sculley, Vijay Janapa Reddi, Lora Aroyo

View PDF

Abstract:The generative AI revolution in recent years has been spurred by an expansion in compute power and data quantity, which together enable extensive pre-training of powerful text-to-image (T2I) models. With their greater capabilities to generate realistic and creative content, these T2I models like DALL-E, MidJourney, Imagen or Stable Diffusion are reaching ever wider audiences. Any unsafe behaviors inherited from pretraining on uncurated internet-scraped datasets thus have the potential to cause wide-reaching harm, for example, through generated images which are violent, sexually explicit, or contain biased and derogatory stereotypes. Despite this risk of harm, we lack systematic and structured evaluation datasets to scrutinize model behavior, especially adversarial attacks that bypass existing safety filters. A typical bottleneck in safety evaluation is achieving a wide coverage of different types of challenging examples in the evaluation set, i.e., identifying 'unknown unknowns' or long-tail problems. To address this need, we introduce the Adversarial Nibbler challenge. The goal of this challenge is to crowdsource a diverse set of failure modes and reward challenge participants for successfully finding safety vulnerabilities in current state-of-the-art T2I models. Ultimately, we aim to provide greater awareness of these issues and assist developers in improving the future safety and reliability of generative AI models. Adversarial Nibbler is a data-centric challenge, part of the DataPerf challenge suite, organized and supported by Kaggle and MLCommons.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
MSC classes:	14J68 (Primary)
Cite as:	arXiv:2305.14384 [cs.LG]
	(or arXiv:2305.14384v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2305.14384

Submission history

From: Lora Aroyo [view email]
[v1] Mon, 22 May 2023 15:02:40 UTC (10,010 KB)

Computer Science > Machine Learning

Title:Adversarial Nibbler: A Data-Centric Challenge for Improving the Safety of Text-to-Image Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Adversarial Nibbler: A Data-Centric Challenge for Improving the Safety of Text-to-Image Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators