Counterfactual Fairness in Text Classification through Robustness

Garg, Sahaj; Perot, Vincent; Limtiaco, Nicole; Taly, Ankur; Chi, Ed H.; Beutel, Alex

Computer Science > Machine Learning

arXiv:1809.10610v1 (cs)

[Submitted on 27 Sep 2018 (this version), latest version 13 Feb 2019 (v2)]

Title:Counterfactual Fairness in Text Classification through Robustness

Authors:Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H. Chi, Alex Beutel

View PDF

Abstract:In this paper, we study counterfactual fairness in text classification, which asks the question: How would the prediction change if the sensitive attribute discussed in the example were something else? We offer a heuristic for measuring this particular form of fairness in text classifiers by substituting individual tokens pertaining to attributes (e.g. sexual orientation, race, and religion), and describe the relationship with other notions, including individual and group fairness. Further, we offer methods, including hard ablation, blindness, and counterfactual logit pairing, for optimizing this counterfactual fairness metric during model training, bridging the robustness literature and the fairness literature. Empirically, counterfactual logit pairing performs as well as hard ablation and blindness to sensitive tokens, but generalizes better to unseen tokens. Interestingly, we find that in practice, the methods do not significantly harm classifier performance, and have varying tradeoffs with group fairness. These approaches, both for measurement and optimization, provide a new path forward for addressing counterfactual fairness issues.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1809.10610 [cs.LG]
	(or arXiv:1809.10610v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1809.10610

Submission history

From: Sahaj Garg [view email]
[v1] Thu, 27 Sep 2018 16:21:39 UTC (221 KB)
[v2] Wed, 13 Feb 2019 19:09:40 UTC (67 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2018-09

Change to browse by:

cs
stat
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Sahaj Garg
Vincent Perot
Nicole Limtiaco
Ankur Taly
Ed H. Chi

…

export BibTeX citation

Computer Science > Machine Learning

Title:Counterfactual Fairness in Text Classification through Robustness

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Counterfactual Fairness in Text Classification through Robustness

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators