Can LLMs Explain Themselves Counterfactually?

Dehghanighobadi, Zahra; Fischer, Asja; Zafar, Muhammad Bilal

Computer Science > Computation and Language

arXiv:2502.18156 (cs)

[Submitted on 25 Feb 2025]

Title:Can LLMs Explain Themselves Counterfactually?

Authors:Zahra Dehghanighobadi, Asja Fischer, Muhammad Bilal Zafar

View PDF HTML (experimental)

Abstract:Explanations are an important tool for gaining insights into the behavior of ML models, calibrating user trust and ensuring regulatory compliance. Past few years have seen a flurry of post-hoc methods for generating model explanations, many of which involve computing model gradients or solving specially designed optimization problems. However, owing to the remarkable reasoning abilities of Large Language Model (LLMs), self-explanation, that is, prompting the model to explain its outputs has recently emerged as a new paradigm. In this work, we study a specific type of self-explanations, self-generated counterfactual explanations (SCEs). We design tests for measuring the efficacy of LLMs in generating SCEs. Analysis over various LLM families, model sizes, temperature settings, and datasets reveals that LLMs sometimes struggle to generate SCEs. Even when they do, their prediction often does not agree with their own counterfactual reasoning.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2502.18156 [cs.CL]
	(or arXiv:2502.18156v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2502.18156

Submission history

From: Muhammad Bilal Zafar [view email]
[v1] Tue, 25 Feb 2025 12:40:41 UTC (59 KB)

Computer Science > Computation and Language

Title:Can LLMs Explain Themselves Counterfactually?

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Can LLMs Explain Themselves Counterfactually?

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators