Structured PREreview of An assessment of normalization and differential expression methods for miRNA-seq analysis using a realistic benchmark dataset
- Published
- DOI
- 10.5281/zenodo.22737224
- License
- CC BY 4.0
- Does the introduction explain the objective of the research presented in the preprint?
- Yes
- It explains the problem properly, namely the need for reliable methods for miRNA analysis, as RNA-seq methods are not able to account for its specific characteristics. It also points out the gaps in prior benchmarking studies, clearly identifying that a benchmark particularly for normalization and differential expression has not been performed before, especially one incorporating recently developed methods. Lastly, it clearly states its aim and overall approach.
- Are the methods well-suited for this research?
- Somewhat appropriate
- 1. Some methodological details are missing. PCA and heatmap figures (figure 1B and 1C) are presented in Results alone, with no explanation in Methods section. So no information on the input data (raw or normalized or log), if it was scaled, the miRNAs selection criteria, or what clustering methods were applied. Due to this, the reproducibility of the paper is affected. 2. The paper frames as evaluating a broad range of normalization strategies but just tests only two types: within-sample scaling (RPM, RPM_total, MAP) and cross-sample (DESeq2,TMM). Other options such as quantile normalization or compositional data methods were not tested. 3. An uneven comparison in the differential expression benchmark, two different input data types used. Most methods use miRNA-level count matrix whereas miRglmm uses isomiR-level count matrix. Still, the authors do not address whether miRglmm’s strong performance is due to the model or input data. 4. The underlying logic for the ground truth is circular. The authors state that DE ground truth is defined by "monotonicity trends derived from RPM values". But RPM is what they recommend for normalization. So ranking other tools against a baseline that uses their preferred method creates a feedback loop that is not validated against any independent benchmark.
- Are the conclusions supported by the data?
- Somewhat supported
- 1. The paper relies almost entirely on monotonicity as the ground-truth signal, across normalization, differential expression and log2FC tissue-specificity direction. While its limitations are acknowledged, its recommendations seem overconfident due to the narrow validation scope and too much reliance on a single criterion. 2. The authors identified pancreatic contamination in Mouse 3, clearly separating on PC2. They state that removing Mouse 3 did not alter the normalization or DE benchmarking results, but it is based only on supplementary figures and not present in the main text. So, it can not be independently verified. 3. In Results, the authors state that rlog-mean had the highest accuracy and precision for mixture proportions and the lowest error for technical replicate variability. But still the paper recommends RPM instead. The justification is that rlog-mean over-corrects the data by flattening variance between random non-replicate sample pairs.. But this relies on supplementary data and lacks quantitative statistical support in the main text. 4. No clear recommendation is made for log2FC estimation. The log2FC accuracy benchmark doesn't have a clear top method. Instead, the results show a sign-dependent split; RPM-FC and edgeR-v2 do better with positive log2FCs, and miRglmm and NBSR do better with negative log2FCs. The authors do not give a default recommendation for users.
- Are the data presentations, including visualizations, well-suited to represent the data?
- Highly appropriate and clear
- The visualizations are well-designed and logical for each analysis goal. These visuals communicate complex data patterns, like normalization trends and differential expression performance, in a clear format. Supplementary figure are a publication norm as well.
- How clearly do the authors discuss, explain, and interpret their findings and potential next steps for the research?
- Somewhat clearly
- Strengths: The authors carefully interpret main findings. Examples like miR-21a and miR-26a are used to show over-correction showing that they do not take trends at face value. Limitations are also thoroughly discussed and notes reliance on one dataset, use of monotonicity as ground truth, arbitrary thresholds, and no power analysis. Weaknesses: No next steps or future directions section. This includes validation on independent mixture data or checking RPM circularity. Key issues like the gap between rlog-mean and RPM performance, and the log2FC split based on sign, are explained, but they are not brought together for a clear recommendation.
- Is the preprint likely to advance academic knowledge?
- Highly likely
- 1. It fills a gap in the literature by using real tissue-mixture instead of simulations or spike-ins. So, it is able to capture biological and technical noise more effectively. 2. The benchmark includes miRNA-specific tools, like miRglmm and NBSR, which were missing from earlier studies. - However due to method gaps, numerical rankings should be interpreted with caution until they are validated by an independent dataset.
- Would it benefit from language editing?
- No
- The paper is well-written and clearly structured with no major issues found.
- Would you recommend this preprint to others?
- Yes, but it needs to be improved
- The authors need to provide methodology details for the figures that are missing. Also, they should evaluate RPM ground-truth circularity and clarify miRglmm's performance. They need to reconcile conflicting metrics such as log2FC and make a clear recommendation.
- Is it ready for attention from an editor, publisher or broader audience?
- Yes, after minor changes
Competing interests
The author declares that they have no competing interests.
Use of Artificial Intelligence (AI)
The author declares that they did not use generative AI to come up with new ideas for their review.