Ir para a Avaliação PREreview

Avalilação PREreview de Tracking claim changes from preprint to publication across 72,644 biomedical studies using large language models

Publicado
DOI
10.5281/zenodo.21907994
Licença
CC BY 4.0

This is a review of version 2 of the paper ‘Tracking claim changes from preprint to publication across 72,644 biomedical studies using large language models’ posted on bioRxiv.

This paper presents a large-scale comparison between the version of an article posted on the bioRxiv preprint server and the subsequent version published in a peer-reviewed journal. Using a large language model, the authors analyze the extent to which the main claims made in an article change between preprint posting and publication in a journal.

I enjoyed reading this interesting paper. I consider the research to be mostly sound and solid. I have two important comments and one smaller comment.

Major comments:

1. “We find that, among the preprints that later reach journal publication, the central claims in abstract largely persist, with about 90% unchanged or minorly revised, while major revisions are observed in 10%. Therefore, bioRxiv preprints may serve as a reliable early source of biomedical knowledge.”: The second sentence doesn’t follow logically from the first one, because the first sentence is only about ‘preprints that later reach journal publication’, while the second sentence is about preprints in general. It could be that preprints that are not published in a journal are of significantly lower quality than preprints that are published in a journal. The authors didn’t study this. Because they didn’t study this, they cannot draw the conclusion that preprints in general (not only those published in a journal) ‘may serve as a reliable early source of biomedical knowledge’.

The limitation that ‘never-published preprints are excluded’ is briefly mentioned in the discussion section. My suggestion would be to discuss this important limitation in a bit more detail.

2. I would like to better understand the performance of the large language model (LLM) for assessing hedging shifts. The statistics reported in Supplementary Figure 2e seem a bit concerning. In the case of human experts, there is a hedging shift for 52 of the 470 pairs (11.1%). In the case of the LLM, however, there is a hedging shift for 120 of the 470 pairs (25.5%). Based on these statistics, the LLM doesn’t seem very successful in reproducing the assessment of hedging shifts by human experts. This then also raises questions about one of the key findings of the paper: “Claims became more cautious twice as often as more confident (8.4% vs 4.2%)”. Are the percentages reported in the paper (8.4% and 4.2%) accurate or not? Also, why does the LLM identify hedging shifts in 25.5% of the cases in Supplementary Figure 2e, but only in 8.4% + 4.2% = 12.6% of the cases in the full analysis?

I suspect the answer to these questions lies in the fact that the statistics reported in Supplementary Figure 2e are based only on pairs for which a majority of the four human experts were in agreement. Pairs for which the human experts did not have sufficient agreement were excluded. To better understand whether the LLM provides accurate assessments of hedging shifts, I would be interested to see a full comparison of the assessments made by the human experts and those made by the LLM.

Minor comment:

In the discussion section, the authors mention the following limitation of their research: “The observed changes (between articles on preprint servers and articles in journals) combine author revision, inputs of reviewers and editors, and journal production.” I think this limitation deserves a bit more reflection. What is not explicitly mentioned by the authors is one of the key benefits of preprinting: By posting an article on a preprint server, researchers make their work available to readers and as a result readers may provide feedback. Researchers in turn may use this feedback to improve their article. This feedback mechanism is not mentioned by the authors. I think it is important and deserves to be acknowledged.

This also means it is not so easy to explain the changes observed between the preprint version of an article and the journal version. I can think of three explanations:

1. Perhaps the most obvious explanation is that changes are the result of the editorial process of the journal to which an article was submitted, and in particular the comments provided by peer reviewers.

2. Changes may result from feedback researchers receive from readers of their preprinted article.

3. Changes may result from researchers making further improvements to their work, independently of the feedback they receive from journal editors, peer reviewers, and others.

Mechanism 1 could be seen as an argument against preprinting. One might argue that prior to peer review articles tend to suffer from quality problems and that peer review is needed to fix these problems. For that reason, preprinting might even be considered undesirable. (Just to be clear, this is not my personal perspective, but this might be the perspective of critics of preprinting.)

Interestingly, mechanism 2 points in the opposite direction. It suggests that preprinting offers a powerful way to get feedback on an article, based on which the article can be improved. This is an argument in support of preprinting.

The paper has the limitation of not being able to distinguish between the above mechanisms. It might be good to offer some further reflection on this limitation.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they did not use generative AI to come up with new ideas for their review.