Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

The Illusion of Explainability with LLMs and LLM-Agents

Authors: Montese, Sara; Alvarez-Napagao, Sergio; Gimenez-Abalos, Victor; Cortés, Ulises;

The Illusion of Explainability with LLMs and LLM-Agents

Abstract

The formal linguistic capabilities of Large Language Models (LLMs) are increasingly intersecting with the field of Explainable Agency (XAg). The growing adoption of LLM-agents has heightened the need to explain their behaviour to users. However, current methods consistently fail to meet desirable properties of human-centric explanations. Moreover, language models are being used to improve traditional agent explainability methods due to their conversational abilities and theirresemblance to human responses, yet are often used without sufficient consideration of their limitations. In this paper, we show how the inclusion of LLM components in the agent architecture affects the reliability of the explanations produced. We argue that the widespread reliance on LLMs in XAg is polluting the definition of explanation and explainability. We question current methods of agentic explainability and argue that their risks may undermine trust. Through an architectural analysis and an empirical illustration, we highlight how certain design choices may limit the kinds of explanatory queries that can be reliably answered. Finally, we propose what human-oriented explainability should entail in LLM-agents, and we expose the limitations and opportunities of LLM s’ integration into agent explainability.

Keywords

Large Language Models, Explainable Artificial Intelligence, LLM-Agents, Agent Explainability

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Funded by