Explanations of Black-Box Models based on Directional Feature Interactions

Masoomi, Aria; Hill, Davin; Xu, Zhonghui; Hersh, Craig P; Silverman, Edwin K.; Castaldi, Peter J.; Ioannidis, Stratis; Dy, Jennifer

Computer Science > Machine Learning

arXiv:2304.07670 (cs)

[Submitted on 16 Apr 2023]

Title:Explanations of Black-Box Models based on Directional Feature Interactions

Authors:Aria Masoomi, Davin Hill, Zhonghui Xu, Craig P Hersh, Edwin K. Silverman, Peter J. Castaldi, Stratis Ioannidis, Jennifer Dy

View PDF

Abstract:As machine learning algorithms are deployed ubiquitously to a variety of domains, it is imperative to make these often black-box models transparent. Several recent works explain black-box models by capturing the most influential features for prediction per instance; such explanation methods are univariate, as they characterize importance per feature. We extend univariate explanation to a higher-order; this enhances explainability, as bivariate methods can capture feature interactions in black-box models, represented as a directed graph. Analyzing this graph enables us to discover groups of features that are equally important (i.e., interchangeable), while the notion of directionality allows us to identify the most influential features. We apply our bivariate method on Shapley value explanations, and experimentally demonstrate the ability of directional explanations to discover feature interactions. We show the superiority of our method against state-of-the-art on CIFAR10, IMDB, Census, Divorce, Drug, and gene data.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2304.07670 [cs.LG]
	(or arXiv:2304.07670v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2304.07670
Journal reference:	International Conference on Learning Representations, 2022

Submission history

From: Davin Hill [view email]
[v1] Sun, 16 Apr 2023 02:00:25 UTC (10,720 KB)

Computer Science > Machine Learning

Title:Explanations of Black-Box Models based on Directional Feature Interactions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Explanations of Black-Box Models based on Directional Feature Interactions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators