Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

Slack, Dylan; Hilgard, Sophie; Jia, Emily; Singh, Sameer; Lakkaraju, Himabindu

Computer Science > Machine Learning

arXiv:1911.02508 (cs)

[Submitted on 6 Nov 2019 (v1), last revised 3 Feb 2020 (this version, v2)]

Title:Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

Authors:Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, Himabindu Lakkaraju

View PDF

Abstract:As machine learning black boxes are increasingly being deployed in domains such as healthcare and criminal justice, there is growing emphasis on building tools and techniques for explaining these black boxes in an interpretable manner. Such explanations are being leveraged by domain experts to diagnose systematic errors and underlying biases of black boxes. In this paper, we demonstrate that post hoc explanations techniques that rely on input perturbations, such as LIME and SHAP, are not reliable. Specifically, we propose a novel scaffolding technique that effectively hides the biases of any given classifier by allowing an adversarial entity to craft an arbitrary desired explanation. Our approach can be used to scaffold any biased classifier in such a way that its predictions on the input data distribution still remain biased, but the post hoc explanations of the scaffolded classifier look innocuous. Using extensive evaluation with multiple real-world datasets (including COMPAS), we demonstrate how extremely biased (racist) classifiers crafted by our framework can easily fool popular explanation techniques such as LIME and SHAP into generating innocuous explanations which do not reflect the underlying biases.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:1911.02508 [cs.LG]
	(or arXiv:1911.02508v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1911.02508

Submission history

From: Dylan Slack [view email]
[v1] Wed, 6 Nov 2019 17:52:20 UTC (962 KB)
[v2] Mon, 3 Feb 2020 18:53:50 UTC (1,115 KB)

Computer Science > Machine Learning

Title:Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators