Skip to main content

Write a comment

PREreview of Demanding peer review is associated with higher impact in published science

Published
DOI
10.5281/zenodo.21515890
License
CC BY 4.0

This work seeks to explain if comprehensive peer review is associated with higher citation impact. The study utilizes LLMs to parse referee and author correspondences, finding that papers with more demanding review (stronger criticism, higher-quality comments, and greater revision cost) tend to result in higher later impact when compared with papers with smoother review processes. This work also finds that referee comments on the novelty/contribution of the work are among the least constructive, and that stronger reviews are also more likely to contain disagreement across reviewers. The findings can help publishers and editors update their peer review policies to promote reviews that focus on the central claims of the manuscripts rather than formatting.

Major issues

  • In Results Section D, the study does not control for discipline when measuring impact. This might reduce the nuance of the results because some fields tend to cite more than others, and social sciences and humanities disciplines tend to receive citations over longer periods of time, usually more than the three years measured in the paper. The paper should explain how disciplinary differences in citations were taken into account or justify why they were not taken into consideration.

  • In the discussion, we recommend acknowledging as a limitation that the results are based only on the Nature Communications journal, which is already among the top journals in several fields (see here). Therefore, the papers published in Nature Communications are already likely to have a higher impact or number of citations (regardless of peer review quality) than smaller or discipline-specific journals. This issue points to the importance of more journals openly sharing their peer review reports so this type of studies can be expanded.

  • We suggest moving Results Section A to the Methodology section. Section A does not provide results to answer the research question, but tells us how the authors constructed and validated the LLM pipeline they developed, so it would be a better fit in the Methods section.

  • Some of the judging criteria for opinion strength, constructiveness, comment quality, and revision cost seem very similar in description. Even humans perhaps would not sort comments into the same numbers 100% of the time; how can we trust an LLM to do so, when it is known that they struggle with nuance? Perhaps it might be better to use fewer than 10 categories that are more clearly separated. If 10 categories are necessary, it would be helpful for the authors to explain why.

Minor issues

  • When describing reviewing disagreement, we suggest switching the adjective “harsh” to “tough” or “challenging” reviews. Harsh could have negative connotations or even imply unprofessional reviews. A different term might help readers avoid getting these connotations.

  • In Results Section A, we suggest that the authors give basic definitions of how opinion strength, constructiveness, comment quality, and revision cost scores are determined. In Section B authors define terms like rebuttal rate, so these scores should also be similarly defined. We realize that the scores are defined in more detail in the methods/supplementary materials, but a basic definition would help guide the reader without requiring them to jump to the supplementary material. In Results Section C, consensus threshold should similarly be defined.

  • In Results Section D, the review-impact association paragraph is unclear. For example, why are the authors multiplying the top and bottom 5% of opinion strength within the same year with opinion-type stratum and response-type stratum? The authors say that the resulting heterogeneity is substantial, but what is the heterogeneity? It would be helpful for this paragraph to be expanded more to clarify these concepts.

  • In the Discussion, it would be useful for the authors to compare and cite other works that discuss similar or related findings. If there are none, it would be useful for the authors to directly say that their work is the first to delve into this kind of analysis. Some papers to consider are:

    • Ross-Hellauer, T., & Horbach, S. P. J. M. (2024). Additional experiments required: A scoping review of recent evidence on key aspects of Open Peer Review. Research Evaluation, 33, rvae004. https://doi.org/10.1093/reseval/rvae004

    • Barnett, A., & Spick, M. (2026). Who chooses open peer review and is it an indicator of article quality? An observational study of PLOS journals. MetArXiv. https://doi.org/10.31222/osf.io/2b7z5_v1

    • Álvarez-García, E., García-Costa, D., Squazzoni, F., Malički, M., Mehmani, B., & Grimaldo, F. (2026). Published peer review reports have higher informative content than unpublished reports. Journal of Informetrics, 20(1), 101760. https://doi.org/10.1016/j.joi.2025.101760

  • The authors should consider addressing the difference between unprofessional reviews and harsh reviews that are still constructive. Can their LLM analysis separate these two cases well? Are unprofessional reviews common/less common in any of the cases they examine?

Competing interests

The authors declare that they have no competing interests.

Use of Artificial Intelligence (AI)

The authors declare that they did not use generative AI to come up with new ideas for their review.

You can write a comment on this PREreview of Demanding peer review is associated with higher impact in published science.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now