Avalilação PREreview Estruturada de Tracing Fairness Disparities Across the Machine Learning Pipeline: A Stage-Resolved Bias Provenance and Propagation Framework for Educational Prediction
- Publicado
- DOI
- 10.5281/zenodo.21823892
- Licença
- CC BY 4.0
- Does the introduction explain the objective of the research presented in the preprint?
- Yes
- Yes. The introduction clearly explains the research objective. It identifies the limitations of end-of-pipeline fairness audits and states that the study aims to develop and evaluate the Bias Provenance and Propagation framework for locating and quantifying disparities across stages of the machine-learning pipeline.
- Are the methods well-suited for this research?
- Somewhat appropriate
- Somewhat appropriate. The methods are generally well suited to the research objectives and include several rigorous elements, such as pre-specified research questions and hypotheses, separate development, validation, and locked test sets, power analysis, bootstrap confidence intervals, multiple-testing correction, and a reference-model sensitivity analysis. However, the robustness and generalisability of the findings are limited by the use of a single dataset, the absence of clickstream data, the dependence of early-stage disparity estimates on the selected reference model, and the non-causal nature of the feature-attribution analysis.
- Are the conclusions supported by the data?
- Somewhat supported
- The conclusions generally reflect the observed stage-level patterns, and the authors appropriately acknowledge that most pre-specified hypotheses were not supported and that all observed disparities were below the minimum effect of interest. However, some of the headline conclusions are stated more definitively than the statistical evidence permits. In particular, claims that disparities originate or are introduced at specific stages rely partly on small, sub-threshold effects whose confidence intervals often include zero. The conclusions should therefore distinguish more consistently between descriptive patterns, statistically supported findings, and hypotheses requiring replication. The use of a single dataset also limits the extent to which the findings can support broader claims about educational machine-learning pipelines.
- Are the data presentations, including visualizations, well-suited to represent the data?
- Neither appropriate and clear nor inappropriate and unclear
- The data presentations communicate the main results reasonably well. The figures include labelled axes, confidence intervals, reference lines, and numerical annotations, while the bias-ledger tables provide the underlying stage-level estimates. However, the visualizations do not fully follow accessibility best practices. Several figures rely primarily on colour to distinguish groups, attributes, or positive and negative contributions, without consistently using alternative line styles, marker shapes, or patterns. The embedded figures also lack descriptive alternative text, and some labels may be difficult to read at reduced size or in grayscale. The authors should improve colour-independent differentiation, add accessible figure descriptions, and ensure that captions communicate the principal numerical patterns independently of the images.
- How clearly do the authors discuss, explain, and interpret their findings and potential next steps for the research?
- Somewhat clearly
- Somewhat clearly. The discussion is generally well organised and provides a substantive interpretation of the main findings. The authors explicitly address the fact that H2, H3, and H4 were not supported and that H1 was supported only for sex/gender. They also distinguish between statistically confirmed effects and descriptive stage-level patterns, noting that all observed disparities were below the pre-specified minimum effect of interest of 0.08. The practical discussion is useful because it connects different disparity profiles with different potential intervention points, such as examining data representation for disability-related disparities and auditing engineered features for possible proxies in the case of sex/gender. The authors also acknowledge several important limitations, including reliance on a single dataset, the absence of VLE clickstream data, the dependence of early-stage disparity estimates on the reference model, the non-causal and model-dependent nature of the SHAP analysis, and the fact that the distribution-shift analysis is simulated rather than based on a real deployment. These limitations indicate relevant directions for future research, particularly replication across multiple educational datasets, evaluation with richer behavioural features, comparison of additional reference models, and validation in real-world deployment and feedback-loop settings. However, the interpretation is not consistently cautious. Some descriptive or statistically unsupported findings are presented using relatively strong language, such as stating that disparities are “manufactured” at feature engineering or treating weak concordance across attributes as a definitive demonstration that no universal origin stage exists. Because the observed effects are small, several confidence intervals include zero, and the evidence comes from one dataset, these statements should be framed more explicitly as dataset-specific and provisional. In addition, the proposed next steps are distributed across the discussion and limitations sections rather than presented as a prioritised research agenda specifying which aspects of the BPP framework require replication, robustness testing, or external validation first.
- Is the preprint likely to advance academic knowledge?
- Moderately likely
- The preprint proposes a useful stage-based framework for tracing fairness disparities across the machine-learning pipeline. However, its contribution is limited by reliance on one dataset, small effects below the minimum effect of interest, unsupported hypotheses, and sensitivity to the reference model. Further replication and external validation are needed.
- Would it benefit from language editing?
- No
- The manuscript is generally clear, coherent, and professionally written. Some wording could be made more cautious or precise, but these issues do not significantly hinder comprehension.
- Would you recommend this preprint to others?
- Yes, but it needs to be improved
- The preprint presents a useful and potentially important framework for stage-resolved fairness auditing. However, some conclusions should be stated more cautiously, and the framework requires broader validation across additional datasets, reference models, and deployment settings.
- Is it ready for attention from an editor, publisher or broader audience?
- Yes, after minor changes
- The preprint is sufficiently coherent, methodologically developed, and clearly presented to merit attention from editors and a broader academic audience. However, some claims should be framed more cautiously, the distinction between confirmed and descriptive findings should be strengthened, and the accessibility of the visualizations should be improved.
Competing interests
The authors declare that they have no competing interests.
Use of Artificial Intelligence (AI)
The authors declare that they did not use generative AI to come up with new ideas for their review.