Avalilação PREreview de Comparative Analysis of Clinical Finding Retrieval Difficulty in Chest X-ray Retrieval Models
- Publicado
- DOI
- 10.5281/zenodo.22788887
- Licença
- CC BY 4.0
In order to evaluate the impact of the candidate-sentence pool on the retrieval of clinical findings for chest X-ray images, we investigate three different sentence-retrieval models, namely the proposed JoImTeRNet model as well as CXR-RePaiR and CXR-ReDonE. We assess these models under two candidate-sentence pool conditions: an original, unbalanced pool of 129,906 sentences and a balanced pool of 6,216 sentences, capped at 450 positive sentences per finding. Our results show that candidate-pool composition strongly affects both finding-level retrieval performance and cross-model agreement on which findings are “hard” to retrieve. Cross-model agreement was strong for the original pool of sentences (ρ=0.749–0.807), but dropped for the balanced pool of sentences (ρ=0.240–0.495). Lung Opacity and Pleural Effusion were consistently the two hardest findings to retrieve, whereas Support Devices was consistently the easiest.
Field contribution
It is novel to empirically demonstrate that the apparent “consensus” between independently trained retrieval models in a multi-model setup does not necessarily stem from the models themselves, but rather may be influenced by the design of the evaluation pool.
Major issues
Confounded variables: The authors note that the “balanced” pool differs from the original pool in terms of size, class balance, and even the included sentences. However, rather than subsampling to a fixed size, the authors use a substantially smaller balanced pool. An equal-size random subsample would be needed to isolate the effect of balancing from the effect of pool size.
Possible bias favoring the proposed model: The sentence pool for evaluation was selected from the MIMIC-CXR database used to train the proposed model, while CXR-ReDonE was trained on the stylistically different MIMIC-PRO. This could confer an advantage unrelated to genuine retrieval quality.
Evaluation labels not expert-verified: The evaluation set is a random sample from the CheXpert training set, not the expert-adjudicated validation set. Moreover, the ground truth was generated by an automated labeler (CheXbert), hence introducing label noise.
Small number of findings (14) for correlation analysis: With Spearman’s rank correlation based on only 14 findings, estimates will inherently be unstable, even with the use of bootstrap confidence intervals.
Speculative explanation for Lung Opacity’s difficulty: The hypothesis in the Figure that the findings’ semantic/definitional ambiguity contributes to the difficulty of retrieving Lung Opacity is plausible but was never directly tested.
Minor issues
Some redundancy between the abstract, results summary, and conclusion.
The “Related Work” section is rather long in comparison to the content of the paper; it could be shortened.
Some of the tables, especially Table 3, contain a lot of numbers and therefore would benefit from a visual summary alongside Figure 1.
Justification for the specific number of 450 positive sentences for the balancing cap is not clear.
Reference [21] is listed as a source but is not used in the paper.
Competing interests
The author declares that they have no competing interests.
Use of Artificial Intelligence (AI)
The author declares that they did not use generative AI to come up with new ideas for their review.