PREreview of Tracking claim changes from preprint to publication across 72,644 biomedical studies using large language models
- Published
- DOI
- 10.5281/zenodo.23046638
- License
- CC BY 4.0
This review is the result of a virtual, collaborative live review discussion organized and hosted by PREreview and the PREreview Club at Future of Research Communication and e-Scholarship (FORCE11) on September 16, 2026. The discussion was joined by 16 people: a facilitator and 13 additional live and two asynchronous review participants. The authors of this review have dedicated additional asynchronous time over the course of two weeks to help compose this final report using the notes from the live and asynchronous reviews. Abeer Abdoon, Amirhosein Sabzian, Morufu Olalekan Raimi, and Xing Jian also participated in the discussion. We thank all participants who contributed to the discussion and made it possible for us to provide feedback on this preprint.
Summary
This study looks at how the scientific claims in bioRxiv preprint abstracts change once they get published as peer-reviewed articles, and what that means for trusting preprints as a source of scientific information. The authors matched 72,644 bioRxiv preprints posted between 2018 and 2025 to their eventual journal-published versions by DOI, then used an LLM (Claude Sonnet 4.6) to extract one primary and two secondary claims from each abstract and classify how much the content, claim type, and certainty shifted between the two versions. A subset of 550 pairs was also checked against four human raters to validate the LLM's classifications.
The main finding reported by the work is that most claims remain stable: 39.8% of primary claims in the preprints were unchanged in the published version, 50.0% underwent only minor revision, and just 10.2% saw major changes, meaning approximately 90% of claims stayed the same or shifted only slightly. Where wordings did change, they tended to become more cautious rather than more confident. Reviewers also flagged the sharp drop in median time between preprint posting and publication in the set of preprints the authors analyzed, from 666 days in 2019 to 160 days in 2024, as an important finding that deserves more attention than the paper currently gives it.
We consistently praised the large scale of the dataset and the fact that the study compares actual claims rather than relying on simple text similarity, reflecting a real methodological improvement over previous works. The study preregistration, open availability of the data, code, prompts, and codebook was also highlighted as a strength.
However, the most frequently raised concern is that the corpus only includes bioRxiv preprints with an identified published paper, which limits how far the findings can be generalized. It was also noted that the study doesn't clearly separate changes introduced through peer review from changes authors made independently, leaving the actual driver of the observed revisions unknown. Additional open questions include the reliability and transparency of the LLM classification process itself, whether the 550-pair human validation sample is large enough to fully support the automated results, and how the overall study population was selected.
Suggestions were made that the work could be improved by widening the scope of the preprint / publication pairs suggested. In addition, the group felt that future work should try to separate the effects of peer review, author-driven revision, time-to-publication, and journal type on how scientific claims evolve.
Evidence and Examples:
With regard to framing and scope, the title implies that whole studies were tracked. However, only preprint abstracts were analysed. For that reason we recommend that the title should be revised. Also importantly, claim stability is not the same as reliability. A claim that survives unchanged does not automatically guarantee that it is correct or robust: limited change in claims may be due to limited inspection during journal-organized peer review rather than robustness of claims. We consequently recommend reframing the paper around the stability and evolution of claims. We also recommend narrowing the conclusions to the stability of abstract claims among published bioRxiv preprints. The headline figure could either be reported against the full bioRxiv corpus or its denominator should be stated explicitly. A discussion of the risks of automated claim assessment by a LLM could also be added. Extending the analysis to full texts would be a valuable future study, but we accept that this is beyond the scope of this work.
The reviewer group also felt that the selection criteria for corpus construction needed clarifying and justifying. The paper analyses around 72K pairs, while the corpus browser appears to hold around 100K. The date windows are also inconsistent: preprints run from 2018 to 2025 and publications from January 2021 to February 2025, yet 101 preprint records in the GitHub file full_corpus_labels.csv were first posted between 2014 and 2017. Early preprints that are published quickly are dropped. Recent manuscripts conversely appear only if they are published fast. This means confoundment, because ‘quick published papers’ are more likely to see claims preserved, and ‘slow published papers’ are more likely to see claims changed. This introduces a strong longitudinal bias.
Each yearly cohort therefore differs for reasons related to inclusion or exclusion, which affects the reported decline in major revisions. The supplementary analysis partly addresses this, but the early-year problem and the corpus definition remain unresolved. The authors could harmonise the date criteria across the repository, manuscript and figures. The authors could also consider modelling fast and slow acceptors explicitly. In addition, some members of the group felt that presenting the numbers as a % of the reduced corpus (i.e. only those that are published) could be misleading. By expressing the numbers as a % of the total corpus, a more accurate estimate of the claims that are made in preprints and are then translated into published form, intact, would be achieved.
The GitHub README and the manuscript describe the inclusion rules only at a high level. The group felt that reproducibility of the corpus could helpfully be improved by adding the following items:
the corpus-construction script
the API retrieval date and query parameters
all filtering and exclusion rules, such as abstract length and language
a record-flow table showing matched and unmatched counts, exclusions at each stage, and the records processed and retained, for instance a table akin to a PRISMA chart.
Because the same model both extracts and compares claims, the LLM classification and validation reporting needs to be fuller. It should explain how the 550 pairs were selected and how the four raters were trained. Agreement should be reported by field and by claim level, not only overall (around 70%; κ 0.63–0.66). Run-to-run variability should also be included. Prompt and version history and data governance details should also be provided. Concise operational definitions of unchanged, minor and major belong in the manuscript, ideally in a table or figure that explains the scoring choices. At present they appear only in the project GitHub repository within prompt_v7.1.md. The detailed decision rules can go in the supplement.
The pairs are treated as independent. In reality however they are highly likely to cluster by journal, field and year (albeit the group did not formally test this). For statistical analysis, mixed-effects models would be more appropriate. At this sample size almost anything could be significant, so effect sizes should take precedence over p-values. Selection bias, confounding and incomplete follow-up also need to be addressed. The data could support stronger tests of association, or of causal effects of peer review, but only once preprint-to-publication time is accounted for.
It was unclear whether the retraction analysis was intended to be exploratory or confirmatory, and its placement within the preprint seemed inconsistent. We recommend that the authors either integrate it fully into the paper, or describe it clearly as exploratory and move it to supplementary material.
The group felt that the structure of the paper needed tightening. Data presentation begins on p. 5 and should be labelled as Results. We were in agreement that the sharing of the labelled corpus and analysis scripts on GitHub was to be commended, and encourage the authors to add the scripts that generate the manuscript's figures.
The shift towards more cautious wording deserves more attention in the opinion of the group. Publication may change how certainty is expressed, not only what is claimed. The design cannot attribute this shift to peer review, editorial intervention, additional analyses or the authors' own revisions. It should therefore be presented as a direction for future research into how different stages and actors in scholarly communication shape the expression of scientific certainty.
Concluding remarks
Overall, the group felt that the preprint was of high quality and a useful contribution to the debate. Because the preprint showed all percentages as a % of preprints that made it to publication, however, this fundamentally overstated the prior expectation of stability that could be applied to the claims in a ‘new’ preprint, and in general the group would like to see more of the findings presented in the context of the overall preprint landscape, not only that part of the preprint landscape that was successfully accepted, as this is a biased sample of preprints that have met some form of quality assessment.
This review represents the opinions of the authors and does not represent the position of Future of Research Communication and e-Scholarship (FORCE11) as an organization.
Competing interests
The authors declare that they have no competing interests.
Use of Artificial Intelligence (AI)
The authors declare that they did not use generative AI to come up with new ideas for their review.