Explainability of Machine Learning Models for Well-Being Prediction

(2026)

Files

STRUMAN_55972000_2026.pdf
  • Open access
  • Adobe PDF
  • 1.8 MB

STRUMAN_55972000_2026_APPENDIX1.py
  • Open access
  • Unknown
  • 19.14 KB

Details

Supervisors
Faculty
Degree label
Abstract
This thesis examines the balance between predictive performance and explainability in machine learning models for well-being prediction. Using questionnaire data from 1,881 individuals, several models including Ridge Regression, CatBoost, LightGBM, and a Stacking Regressor were evaluated with cross-validation. Results show that the Stacking model achieved the lowest prediction error, but its advantage over other strong models, including the interpretable Ridge model, was limited and not statistically significant. To assess interpretability, SHAP and LIME were applied to analyze feature importance. At the global level, models relied on similar key predictors, while local explanations were more model-dependent. Additional analyses confirmed that explanations were generally consistent, stable, and aligned with model behavior. Overall, black-box models offer only modest performance gains and remain less transparent than linear models. Their use in well-being applications should therefore be justified not only by accuracy, but also by the reliability and clarity of their explanations.