Large Language Models (LLMs) now write with a fluency and persuasiveness that can subtly steer users' choices. When their outputs lack clear and comprehensible explanations, this persuasive power risks undermining human decision-making capacity, raising serious ethical concerns. Current explainable artificial intelligence (XAI) techniques focus primarily on technical transparency for epistemic purposes (how a model works); they are rarely intended to reveal to the user the kind of influence they are subject to. Drawing on the Indifference View of manipulation, we advance a preliminary framework that reconceives explainability as both an epistemic and an ethical imperative. The core idea is based on explanatory metadata: layered annotations that accompany model outputs with four complementary types of explanation-informative, justificatory, causal, and precautionary-which give models the ability to detail the reasons underlying the influence they exert. Doing so shifts the XAI goal from mere transparency to responsible influence. It positions explanations as a safeguard against the manipulative behavior of generative AI systems, laying the groundwork for future methods that measure, audit, and actively constrain ethically problematic influence.

Beyond Technical Transparency: Explainability as a Safeguard Against Manipulative AI*

Petrocchi, Ermanno;Tiribelli, Simona;Calvaresi, Davide
2025-01-01

Abstract

Large Language Models (LLMs) now write with a fluency and persuasiveness that can subtly steer users' choices. When their outputs lack clear and comprehensible explanations, this persuasive power risks undermining human decision-making capacity, raising serious ethical concerns. Current explainable artificial intelligence (XAI) techniques focus primarily on technical transparency for epistemic purposes (how a model works); they are rarely intended to reveal to the user the kind of influence they are subject to. Drawing on the Indifference View of manipulation, we advance a preliminary framework that reconceives explainability as both an epistemic and an ethical imperative. The core idea is based on explanatory metadata: layered annotations that accompany model outputs with four complementary types of explanation-informative, justificatory, causal, and precautionary-which give models the ability to detail the reasons underlying the influence they exert. Doing so shifts the XAI goal from mere transparency to responsible influence. It positions explanations as a safeguard against the manipulative behavior of generative AI systems, laying the groundwork for future methods that measure, audit, and actively constrain ethically problematic influence.
2025
File in questo prodotto:
File Dimensione Formato  
2025 Beyond Technical Transparency. Explainability as a Safeguard Against Manipulative AI.pdf

solo utenti autorizzati

Descrizione: Article
Tipologia: Documento in post-print (versione successiva alla peer review e accettata per la pubblicazione)
Licenza: Copyright dell'editore
Dimensione 331.63 kB
Formato Adobe PDF
331.63 kB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11393/382030
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact