Aller directement au contenu principal

Rédiger un PREreview

Tracing Fairness Disparities Across the Machine Learning Pipeline: A Stage-Resolved Bias Provenance and Propagation Framework for Educational Prediction

Publié
Serveur de preprints
EdArXiv
DOI
10.35542/osf.io/dhmwb_v1

Predictive models are now embedded across education, flagging at-risk students and shaping the support they receive. Most fairness audits measure group disparity in the final prediction and then apply a fix, but this end-of-pipeline view cannot say where the unfairness originated or how far it travelled. We introduce Bias Provenance and Propagation (BPP), a measurement framework—not a new mitigation method—that treats disparity as a quantity with a source stage and a path through the machine-learning pipeline. Five stages (sampling, preprocessing, feature engineering, model training, evaluation) plus an exploratory simulation of distribution shift are placed on one threshold-independent scale, the Absolute Between-ROC Area (ABROCA), by projecting each stage's information through a single frozen reference model. Two derived measures form a bias ledger that records where fairness debt is incurred and repaid, each with a bootstrap confidence interval. Applied to the Open University Learning Analytics Dataset (32,593 records) on a leakage-free early-warning task, BPP shows that the origin of disparity is attribute-specific: for disability it is present in the raw data and attenuates once early assessment and registration features enter, whereas for sex/gender and deprivation it is introduced at feature engineering. Several pre-specified hypotheses were not supported—in particular, no universal origin stage exists (Kendall's W = 0.11)—and, tellingly, output-only post-processing reduces the operating-point gap while leaving ABROCA untouched. All observed disparities were modest, below our pre-specified 0.08 minimum effect of interest, so we frame them as consistent, stage-localised patterns rather than confirmed effects. BPP reframes fairness auditing as an accounting problem across stages, showing practitioners where a fix will actually pay off.

Vous pouvez rédiger un PREreview de Tracing Fairness Disparities Across the Machine Learning Pipeline: A Stage-Resolved Bias Provenance and Propagation Framework for Educational Prediction. Un PREreview est une évaluation d'un preprint et peut varier de quelques phrases à un rapport détaillé, semblable à un rapport d'évaluation par les pairs organisé par une revue.

Avant de commencer

Nous vous demanderons de vous connecter avec votre identifiant ORCID iD. Si vous n'en avez pas, vous pouvez en créer un.

Qu’est-ce qu’un ORCID iD ?

Un ORCID iD est un identifiant unique qui vous distingue de toute personne ayant le même nom ou nom similaire.

Commencer maintenant