Evaluating the Effectiveness of Resampling Strategies to Improve Fairness in Machine Learning

Loading...
Thumbnail Image

Files

Jayasekera_umd_0117E_25900.pdf (1.86 MB)
(RESTRICTED ACCESS)
No. of downloads:

Publication or External Link

External Link to Data Files

Date

Advisor

Sweet, Tracy M
Stapleton, Laura M

Citation

Abstract

This study explored the performance of various data resampling methods when used in conjunction with a variety of machine learning (ML) algorithms to produce fair predictions. These methods and algorithms were applied to data that were generated to have various levels of outcome class imbalance and sensitive predictor group imbalance. Data were also generated to be intentionally biased through the introduction of moderators based on sensitive predictor group membership to induce differential (or heterogeneous) predictor effects. This dissertation is highly relevant in the current landscape where machine learning models are increasingly used to make critical decisions across various sectors, from healthcare and education to criminal justice. As these algorithms often rely on data with distributional imbalances and/or differential predictor effects, it is imperative to address these challenges to ensure fairness and accuracy in predictions, especially when sensitive predictors (e.g., race, gender, socioeconomic status) are involved.

The gap that this study aimed to fill was significant. Although much of the existing research focused on how to handle outcome class imbalances, few studies have explored how preprocessing methods, which were developed to handle imbalanced outcome classes, can address fairness in predictive scenarios when there are imbalanced sensitive predictor groups. The distinction was crucial because the objective here was not just to balance outcomes so that algorithms were better able to predict minority outcome classes, but to create fairness in the predictions themselves, ensuring that the predictions made by the algorithms were fair across different demographic subgroups, which could have a much broader societal impact.

Addressing the gap in the literature, this study examined the suitability of methods traditionally used to handle outcome class imbalance, such as random oversampling and SMOTE, in the context of fairness. It investigated when and how these methods may outperform algorithmic predictions made without any data resampling. To this end, the study first employed a Monte Carlo simulation to evaluate the predictive performance of machine learning algorithms combined with oversampling methods and techniques under various conditions designed to emulate the complexities of real-world data. A secondary study investigated the efficacy of these methods compared to the separate application of algorithms based on sensitive group membership. Finally, an empirical analysis was conducted to demonstrate the effectiveness of resampling methods in practice, focusing on predicting five-year degree attainment accurately and fairly across student subpopulations using a large-scale education dataset from the Maryland Longitudinal Data System (MLDS).

Notes

Rights