Comentarios
Escribir un comentarioNo se han publicado comentarios aún.
The research question is relevant and important, particularly because clinical AI tools can directly influence patient care and healthcare decision-making. The study addresses whether a structured rubric can practically and reproducibly assess ethical issues in clinical AI tools before and after deployment.
Strengths
1. The rubric provides a concise and practical framework that could be useful for clinicians and other healthcare professionals who are not AI specialists.
2. The five domains—Evaluation Population, Accessibility, Clinical Bias, User Control, and Monitoring and Updates—cover several important ethical dimensions.
3. Using publicly available information makes the assessment more realistic for hospital procurement boards and clinical departments that may not have access to proprietary AI-development data.
4. The study has a multidisciplinary perspective involving clinical and AI expertise.
5. The tables provide a relatively clear presentation of the scoring framework.
Scoring methodology
1. A major concern is the use of an unweighted linear composite score.
2. High scores in some domains could potentially compensate for a serious weakness in Clinical Bias or another high-risk domain.
3. The authors should consider explaining why all domains receive equal weight and whether some ethical risks should have greater influence on the overall assessment.
Inter-rater reliability
1. The rubric-development process would be stronger if the authors reported measures of inter-rater reliability, such as Cohen’s kappa or another appropriate statistic.
2. This is particularly important because reproducibility is one of the stated purposes of the rubric.
Publicly available information
1. Restricting the assessment to publicly available information is understandable and has practical value, but it may also result in incomplete assessment of AI tools.
2. The manuscript should clearly acknowledge how missing or unavailable information affects the scoring.
Selection of AI tools
1. Only three AI tools were evaluated.
2. The authors should provide clearer criteria for why these three tools were selected and acknowledge that the small sample limits the generalizability of the findings.
Validation
1. The current study demonstrates the application of the rubric but does not constitute full validation.
2. Further testing with a larger and more diverse set of AI tools and independent reviewers would strengthen evidence for the rubric’s reliability and validity.
Reproducibility and thresholds
1. Table 2 makes the scoring framework relatively clear, but the manuscript does not appear to establish explicit thresholds for what constitutes acceptable, concerning, or unacceptable ethical risk.
2. Defining such thresholds could make the rubric more actionable for healthcare organizations.
Comparison of AI tools
1. The numerical scores (Abridge 8/10, UTICalc 6/10, and Medihub Prostate 3/10) are useful for illustrating the rubric.
2. A visual comparison, such as a bar chart, could make the differences between the tools easier to interpret.
Conclusions
The conclusions generally follow from the findings.
However, statements suggesting that the rubric will improve patient trust, clinical practice, or patient outcomes should be presented as potential benefits rather than established effects, because these outcomes were not directly evaluated in the study.
Limitations
The authors should emphasize the limitations associated with:
the small number of AI tools assessed;
reliance on publicly available information;
absence of full validation; and the focus on clinical AI tools rather than potentially broader applications.
Readability
Some sections could be simplified to improve accessibility, particularly for researchers and reviewers whose first language is not English.
Future research
1. Future studies could test the rubric across a larger and more diverse range of clinical AI tools and healthcare settings.
2. Further research could also examine how the rubric performs when incorporated into actual health-system governance and procurement processes.
Overall assessment
1. The rubric addresses a meaningful gap by providing a structured approach to evaluating ethical considerations in clinical AI tools.
2. Its practical value is promising, but additional validation, clearer risk thresholds, reliability testing, and broader application would strengthen the methodology and support wider adoption.
3. These comments are based directly on the reviewers’ notes in the uploaded Live Review document.
Feedback for Authors:
Abeer Abdoon:
This study addresses an important and timely issue by providing a practical framework for assessing ethical considerations in clinical AI tools. The five-domain rubric is clearly structured and has potential value for clinicians and healthcare organizations, particularly when detailed information about proprietary AI systems is not readily available. The multidisciplinary approach and practical application to clinical AI tools are also strengths.
The manuscript could be strengthened by providing additional justification for the equal weighting of the five domains and by reporting inter-rater reliability to support the reproducibility of the rubric. It would also be helpful to clarify how missing publicly available information affects scoring, provide clearer criteria for selecting the evaluated AI tools, and define thresholds for interpreting overall ethical risk. The authors should be cautious about presenting potential effects on patient trust, clinical practice, or outcomes as established benefits, since these were not directly evaluated. Overall, broader testing with more AI tools and independent reviewers would provide stronger evidence for the rubric's reliability, validity, and applicability.
Nusiba Dafalla
This preprint aims to study the matter of healthcare organizations conducting a systematic, quantitative, and reproducible assessment of the ethical implications of clinical AI tools at pre- and post-deployment phases due to the expanded usage of AI in healthcare field and that is the main contribution of this study.
The clinical AI’s influence on patient care and healthcare decisions is critical, that’s why the main goal of this preprint is to create a scoring system that health care healthcare organizations use to identify and evaluate ethical issues related with AI tools both before and after their deployment. This is important because these clinical AI can affect patients, and healthcare decisions if not been dealt with carefully, and that is an excellent contribution to the clinical field.
What is more fascinating is that trial of the effective usage of publicly available information reflects the constraints faced by hospital procurement boards and clinical departments that may not have access to vendors' own algorithms. This standard also complements existing clinical reporting guidelines by providing a quantitative approach to ethical governance, and giving each of the AI tools they used a number from 1 to 10.
One of the greatest features of this study is including even the researchers who are not familiar with AI usage of knowing how to use it, and also represents a practical checklist that enables non-AI-specialist clinicians to conduct standardized ethical screening. also the details provided in the preprint were enough information to understand the main method and how the rubric was applied.
I would recommend for the author to figure a way to include researchers who are not related to the medical field to use it also, and to deal with the unweighted linear scoring permits strong scores in bureaucratic sections (like Accessibility) to conceal grave risks in essential safety domains (such as Clinical Bias), thus showing a bias towards clinical usage exclusively. Also as for the graphs used (tables) i think it would have been easier if the Authors used bar charts for the comparison. Regardless of that it provides a clear, actionable checklist guidance matrix.
Although the approach used is appropriate for developing and demonstrating a rubric; but has some major issues such as not being sufficient yet for full validation, also the development process relied on expert consensus, but no inter-rater agreement metrics (e.g., Cohen’s kappa) were measured across independent scorers, Summing scores across domains assumes equal ethical weight. A tool with known clinical bias (-1) could still receive an overall positive score (4/10) if other domains score well, also the study relied only on publicly available data this creates limitations.
Nasteha Abubakar
With the Increase use of AI in clinical setting their ethical Implications has to be considered and since these tools influence patient-care,Clinical decision-making its important to have structured approach to identify potential ethical risks before and after their deployment.
The study developed a structured scoring rubric designed to assess ethical issues associate with these tools,the rubric provided a framework that can be used to examine different ethical dimensions,the authors later tested this rubric using three AI tools that are currently used in clinical settings. Overall the proposed scoring system can support healthcare organizations identify and evaluate ethical concerns both before and after the deployment of these AI tools.
The authors declare that they have no competing interests.
The authors declare that they did not use generative AI to come up with new ideas for their review.
No se han publicado comentarios aún.