Skip to main content

Write a comment

PREreview of Comparative Analysis of Explainable AI for Depression Risk Assessment Based on Digital Behavior of University Students

Published
DOI
10.5281/zenodo.23110671
License
CC BY 4.0

1. Introduction and research objectives

We find the research question relevant, particularly the attempt to make predictions about students’ mental health more interpretable. Comparing logistic regression with a random forest is a sensible starting point: it allows the authors to examine whether a more flexible model offers a useful advantage over a simpler baseline.

The intended application is less clear. The introduction discusses continuous monitoring and the limitations of self-report, yet the predictors come from a questionnaire. What would this model add to administering the PHQ-9 directly? Clarifying this point would help readers understand the practical problem the study is intended to solve.

The title also promises a comparison of explainable AI approaches, while the analysis compares two classifiers using the same explanation framework. A narrower title would better reflect the work presented.

2. Suitability and reporting of the methods

Our main concern is the limited information available about the data and the analytical procedure. The choice of algorithms is reasonable for an exploratory study, but the present description does not allow the results to be reproduced or their reliability to be assessed adequately.

For the 94 respondents, we would need to know how participants were recruited, the eligibility criteria, the sample characteristics, and the number of participants in each outcome category. The exact training and test counts, handling of missing or duplicate responses, questionnaire items, and coding scheme should be reported. The meaning and direction of the scales for “academic impact” and “digital dependency” are especially important.

A single 80:20 split is a fragile basis for comparing models in this sample. Repeated stratified cross-validation, if the class counts allow it, would give a better indication of how dependent the findings are on the partition. Both models should use the same splits. Any tuning and preprocessing must be confined to the appropriate training data, with model selection separated from final evaluation.

The manuscript also needs the model settings, software versions, random seeds, and an account of ethical approval or exemption, consent, and data protection. We raise the ethical reporting issue because these details are needed to assess the study; their absence from the text does not establish that safeguards were absent in practice.

3. Support for the conclusions

The reported accuracy needs to be put into perspective. If the test set contained 19 participants, the results correspond to 18 correct classifications for logistic regression and 19 for the random forest. The difference would therefore rest on one case.

Under that assumption, the exact two-sided 95% binomial confidence interval for 19 correct classifications out of 19 is approximately 82.4–100%. This is an illustrative calculation based on an inferred test-set size, not an interval reported in the manuscript. It also leaves out uncertainty associated with model selection and recruitment.

We therefore find the claim of superior random-forest performance insufficiently supported. A perfect result on this test set neither establishes overfitting nor demonstrates a reliable advantage in capturing nonlinear relationships.

The clinical interpretation needs similar restraint. The outcome is a category derived from a PHQ-9 threshold. The analysis does not establish a clinical diagnosis or predict future onset of depression. Its conclusions should remain tied to the questionnaire outcome actually measured.

4. Presentation of results and visualisations

The SHAP plots are a useful inclusion. However, one apparent inconsistency needs to be resolved: in Figure 1, higher encoded values of “digital dependency” have negative SHAP values, whereas the text associates greater dependency with a positive contribution to predicted risk. Reverse coding could explain this. Without the coding scheme and the identity of the explained class, the reader cannot tell whether the problem lies in the wording or in the interpretation.

Table 1 is also incomplete relative to the stated evaluation plan. Precision, recall, and F1-score should be reported alongside confusion matrices, class counts, and uncertainty estimates. Sensitivity, specificity, and a majority-class baseline would make the practical meaning of the results clearer. If the model is intended to provide individual probabilities, calibration should be examined as well.

For the SHAP analysis, the explainer, reference data, explained observations, and output scale need to be specified. The authors should also explain how Table 2’s feature-importance scores were obtained and how they differ from the SHAP summaries.

5. Interpretation of findings and future research

We appreciate the acknowledgement of the small sample and possible overfitting, but the discussion needs to engage more closely with what the models actually show. The differing feature rankings are particularly relevant. Are those rankings stable across resampling, and could correlations between predictors help explain the disagreement?

The dominance of “academic impact” in the random forest deserves closer examination. Does this variable capture perceived effects of technology use on studying, or academic difficulties more generally? Comparing the full model with a model using only this variable, and with a model excluding it, would help establish what the remaining predictors contribute.

The authors should confirm that no predictor was derived from PHQ-9 scores or outcome labels. Overlap in the content of questionnaire items, including sleep-related questions, should also be discussed. Such overlap is not automatically data leakage, but it matters when judging whether the model supplies independent information.

SHAP explanations concern the model’s predictions. They do not establish that the highlighted behaviours cause depressive symptoms or that modifying those behaviours would improve mental health. In our view, stronger internal evaluation should come first, followed by independent external and prospective studies before practical screening or monitoring claims are made.

6. Contribution to academic knowledge

The work has potential as a pilot study. Including a simple baseline and examining explanations for both models are useful choices. At present, however, the contribution is mainly exploratory: applying established classifiers with SHAP does not itself amount to a methodological advance.

We would encourage the authors to state more clearly what this particular sample, set of measurements, or comparison adds to the literature. A reproducible analysis would make that contribution easier to judge. The TRIPOD+AI checklist could help identify missing reporting details, although completing it would not resolve the limitations of the design.

7. Language and editorial improvements

The manuscript is readable. The more pressing editorial task is to make the terminology and claims precise. “Depression,” “depression risk,” and a PHQ-9-defined symptom category should not be used interchangeably. The predictors should consistently be described as self-reported measures, and statements about model superiority should reflect the uncertainty of the comparison.

The references need checking against the claims they support. In the introduction, reference [3], the foundational SHAP paper, is used to support an empirical association between digital behaviour and depressive symptoms. An appropriate empirical source is needed there.

8. Recommendation to other readers

We would recommend reading this preprint as an exploratory contribution to research on student mental health and explainable machine learning. Its main interest lies in the question it raises and the analyses that could follow from it.

Readers should be cautious about treating the reported accuracy as evidence of clinical performance. The study does not yet provide a sufficient basis for institutional screening or decisions about individual students.

9. Readiness for wider consideration

Our assessment is that the manuscript needs substantial revision before its broader conclusions can be supported. The immediate priorities are a fuller account of the measurements and analysis, a more reliable assessment of model performance, complete reporting of the evaluation metrics, and a resolution of the SHAP interpretation issue.

These are substantive concerns, but they also give the authors a clear route for improving the paper. With appropriate reanalysis and more restrained conclusions, the study could become a useful and transparent pilot investigation.

Competing interests

The authors declare that they have no competing interests.

Use of Artificial Intelligence (AI)

The authors declare that they did not use generative AI to come up with new ideas for their review.

You can write a comment on this PREreview of Comparative Analysis of Explainable AI for Depression Risk Assessment Based on Digital Behavior of University Students.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now