Thwarting finite difference adversarial attacks with output randomization

Khan, Haidar; Park, Daniel; Khan, Azer; Yener, Bülent

Computer Science > Machine Learning

arXiv:1905.09871 (cs)

[Submitted on 23 May 2019]

Title:Thwarting finite difference adversarial attacks with output randomization

Authors:Haidar Khan, Daniel Park, Azer Khan, Bülent Yener

View PDF

Abstract:Adversarial examples pose a threat to deep neural network models in a variety of scenarios, from settings where the adversary has complete knowledge of the model and to the opposite "black box" setting. Black box attacks are particularly threatening as the adversary only needs access to the input and output of the model. Defending against black box adversarial example generation attacks is paramount as currently proposed defenses are not effective. Since these types of attacks rely on repeated queries to the model to estimate gradients over input dimensions, we investigate the use of randomization to thwart such adversaries from successfully creating adversarial examples. Randomization applied to the output of the deep neural network model has the potential to confuse potential attackers, however this introduces a tradeoff between accuracy and robustness. We show that for certain types of randomization, we can bound the probability of introducing errors by carefully setting distributional parameters. For the particular case of finite difference black box attacks, we quantify the error introduced by the defense in the finite difference estimate of the gradient. Lastly, we show empirically that the defense can thwart two adaptive black box adversarial attack algorithms.

Subjects:	Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
Cite as:	arXiv:1905.09871 [cs.LG]
	(or arXiv:1905.09871v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1905.09871

Submission history

From: Haidar Khan [view email]
[v1] Thu, 23 May 2019 18:58:39 UTC (805 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2019-05

Change to browse by:

cs
cs.CR
cs.CV
stat
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Haidar Khan
Daniel Park
Azer Khan
Bülent Yener

export BibTeX citation

Computer Science > Machine Learning

Title:Thwarting finite difference adversarial attacks with output randomization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Thwarting finite difference adversarial attacks with output randomization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators