Comparative Analysis of Clinical Finding Retrieval Difficulty in Chest X-ray Retrieval Models
- Posted
- Server
- Research Square
- DOI
- 10.21203/rs.3.rs-11003136/v1
Automated chest X-ray (CXR) report generation increasingly relies on retrieval-based systems that select sentences from fixed candidate pools of clinical language. Because these pools may be unevenly distributed across clinical findings, this study investigated whether candidate-pool composition affects finding-level retrieval performance and cross-model agreement on relative finding difficulty. Three sentence-level image-to-text retrieval models, a proposed model based on the JoImTeRNet framework, CXR-RePaiR, and CXR-ReDonE, were evaluated using finding-level recall, precision, and F1 at K=1, 5, and 10 under two candidate-pool conditions. The original pool contained 129,906 sentences from the MIMIC-CXR findings sections, while a balanced, reduced-size candidate pool capped positive sentences per finding at 450, yielding 6,216 sentences. Cross-model agreement in finding difficulty was assessed using Spearman rank correlation with paired bootstrap testing. Overall recall ranged from 0.09 to 0.31 and precision from 0.15 to 0.29 across models and K values. Lung Opacity and Pleural Effusion were consistently among the most difficult findings to retrieve, whereas Support Devices was consistently among the easiest. Under the original pool, cross-model agreement on finding difficulty was strong (ρ=0.749–0.807) but decreased markedly under the balanced pool (ρ=0.240–0.495), with significant reductions for all model pairs (∆ρ=0.298–0.487, all p<0.001). These findings indicate that candidate-pool composition substantially influences both finding-level retrieval performance and cross-model agreement, indicating that apparent consensus on retrieval difficulty may partly reflect evaluation-pool composition rather than intrinsic model behavior.