Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Conference object
Data sources: ZENODO
addClaim

Explaining Agents' Interactions through their Causal Behavior and Counterfactuals

Authors: Islam, Mir Riyanul; Barua, Shaibal; Ahmed, Mobyen Uddin; Begum, Shahina;

Explaining Agents' Interactions through their Causal Behavior and Counterfactuals

Abstract

Reinforcement learning (RL) agents often operate as black boxes, making it difficult to understand their decision-making in dynamic environments. This study proposes a novel framework for explainable RL based on structural causal models (SCMs). Here, the approach learns an SCM of the environment dynamics and reward process in a mobile network simulator (mobile-env), and uses this causal model to generate counterfactual explanations and perform interventions to understand agent behavior. The approach demonstrates that the learned SCM can closely approximate the environment’s transition dynamics while remaining interpretable. By leveraging do-calculus and counterfactual reasoning, our framework explains the long-term effects of actions through causal chains and highlights key influential factors. Experiments on a wireless network control task show that our method provides meaningful explanations for agent decisions (e.g., why a given action yields a higher reward), with minimal loss in policy performance. The study also presents comparative evaluations against baseline explanation approaches and discusses how our SCM-based explanations improve transparency and trust in RL policies.

Powered by OpenAIRE graph
Found an issue? Give us feedback
Funded by