Robustness of Fairness-Constrained Classification under Data Poisoning Attacks
Files
JULIORUIZGUINAZU_03061900_2026.pdf
Open access - Adobe PDF
- 2.09 MB
Details
- Supervisors
- Faculty
- Degree label
- Abstract
- As automated decision-making systems become increasingly deployed in high- stakes domains, ensuring that fairness-aware classifiers remain robust to adversarial manipulation has become an important challenge. This thesis studies the robustness of the maximum-entropy logistic regression framework proposed by Vancompernolle et al. under fairness-targeted data poisoning attacks, comparing its performance against the decision-boundary covariance model of Zafar et al. and an unconstrained logistic regression baseline. Two attack frameworks are considered: a gradient-based poisoning attack and an optimal label-flipping attack. The three models are evaluated on five real-world benchmark datasets under varying poisoning settings and fairness constraints. A key finding is that both attacks produce no statistically significant effect on the unconstrained baseline, establishing that the vulnerability to poisoning is not an inherent property of logistic regression but is introduced directly by the fairness enforcement mechanism. Among the two fairness-constrained formulations, the experimental results reveal important differences in how this vulnerability manifests. The maximum-entropy approach generally preserves lower demographic parity degradation under attack, while the Zafar formulation appears more vulnerable to fairness deterioration despite maintaining stable predictive accuracy, a silent failure mode that standard performance monitoring would not detect.