Explainability and machine learning (for text)

(2025)

Files

Loncour_25862000_2025.pdf
  • Open access
  • Adobe PDF
  • 4.25 MB

Details

Supervisors
Faculty
Degree label
Abstract
Neural networks such as BERT and GPT-3 have achieved impressive performance across various natural language processing tasks, yet the reasoning behind their decisions remains opaque. As these models are increasingly deployed in sensitive domains, the need for transparent and trustworthy explanations becomes critical. This Master's thesis investigates whether these explanations are stable, meaning that different models trained under the same conditions generate similar explanations. In the first part, using synthetic datasets, this work observes that the length and the coherence of an input text impact that stability and that text that do not contain discriminant markers produce unstable explanations. In the second part, this work concludes that the difficulty of the task is decoupled from the stability of its explanations using real-world classification settings. Through this, the thesis contributes to the broader question of whether complex model decisions can be reliably explained using simple, human-understandable representations. Future research could look into the impact of the number of strong markers in a text on the explanation stability.