Comentarios
Escribir un comentarioNo se han publicado comentarios aún.
I have reviewed this article for a journal and am posting the review here as a public contribution.
Summary:
This is a simple and objective manuscript leveraging LLMs to address the important question of how much content changes between preprints and published versions. This topic has been addressed in specific fields of science in the past, but this analysis has much greater scope, although it is subject to the limitations inherent to automated content analysis. Data and code are available in GitHub and are well documented.
Major points:
1. My main concern with the paper is the fact that, as far as I could understand from the Methods and what appears to have been the prompt (https://github.com/rustlab1/PreprintPaperTracker/blob/main/analysis/codebook/prompt_v7.1.md), claim extraction was performed on each pair of abstracts. If I understood correctly, the model received both abstracts and was told to extract the primary claim and up to two secondary claims based on the pair.
If I’m correct, it seems to me that this procedure inherently forces the model to extract something that is present on both sides – i.e. if the claim changes completely from one version to the other, or if a claim appears or disappears, it seems unlikely that this will be considered a primary or secondary claim for the pair of abstracts.
My worry is that this fact may have biased the subsequent analyses towards greater agreement, as everything that comes downstream effectively works on claims that were selected based on an abstract pair. For instance, the authors do consider the possibility of “primary claim replaced” as one of the forms of major change later. But if the claim is being extracted for the pair of articles, would a claim that was present only one version, would that claim even make it to that part of the analysis?
A reasonable test to see if this bias may be present would be to have a model (or human rater) extract the primary and secondary claims independently in both versions (e.g. blind to the other version) and check whether this affects the claims extracted (and the subsequent analysis). This may not have to be done for the whole sample, but it seems like an important sanity check to ensure that the concordance between abstracts is not being artificially inflated by the extraction procedure.
2. Closely related to the point above, it wasn't clear to me whether the abstract pairs for which claim replacement was detected in the LLM analysis was registered by the model (beyond merely classifying these as major revision). If this was the case, it would be useful to know how frequent this was (note that, from the Methods, it seems to have happened in around 7% of claims extracted by humans), and those cases should probably be removed from the hedging/claim type transition analyses that come later, as these questions do not make sense if the same claim is not present on both sides. If it was not the case, this should be mentioned as a limitation, as it will like bias the remaining analyses.
3. Stil somewhat related to the abovementioned points is the fact that LLM validation results (in terms of agreement with human raters) are provided for the rating questions (i.e. claim change and hedging). Nowhere in the paper, however, I could find a measure of how much they agree in terms of claim extraction (unless the kappa and alpha measures included in the results also include this, but my feeling was that they referred to claim change only, as claim extraction is likely not as amenable to straightforward categorization). Nevertheless, the description in the Methods seems to imply that the whole abstract evaluation process (including claim extraction) was performed by the human raters as well. If this was the case, can the authors provide an estimate of how consistent was the claim extraction process between LLMs and humans?
4. There are a few clarity issues that, although minor, make a very big difference in interpretation and should be addressed. The main one is that nowhere in the title (and only late in the abstract) do we get to learn that the conclusions are based strictly on abstracts. In fact, even in the results this is not clear (“parsed the preprint and published abstract into one primary and two secondary claims” seem to imply that the whole preprint was used, as “abstract” is used as singular and not plural). Eventually (i.e. in the Methods) it becomes clear that both versions are evaluated based on the abstract only, but this is an important issue to be fixed. The title should probably be changed to something like “…72,644 biomedical abstracts”, the abstract should probably read “72,644 preprint abstracts to their peer-reviewed versions”, and the passage from the results quoted above should read “parsed the preprint and published abstracts” in the plural form.
5. This is not exactly a concern, but the fact that the rate of substantial revision declines markedly over time (from 17% to 5.7%) was to me the most intriguing finding of the paper. Do the authors have any hypothesis for why this may have happened? There seems to be no discussion of this finding either in the results or the discussion, and I do think it deserves a bit of reflection.
Minor points:
Abstract:
1. There is a contradiction between what is described here (“extract one primary claim and two secondary claims”) and what is quoted in the Methods (“extract one primary claim and up to two secondary claims”). As far as I could understand from the GitHub prompt the abstract version is the correct one, but please revise and correct where necessary.
2. This is minor, but I’d change “major revisions” in the second-to-last sentence to “substantial revision” to (a) make clear that this corresponds to the same category as that mentioned in the previous sentence and (b) avoid confusion with “major revision” as an editorial decision category.
Introduction:
1. The 65 to 70% figure quoted for rate of preprint publication is from quite a few years back (and belongs to a pre-LLM era), so it may be important to make this timing clear and use the past tense when citing these numbers.
2. I’d argue that the initial rise in preprint usage in biology precedes the Covid-19 pandemic for basic biology (with a turning point in 2016, with developments such as the creation of ASAPbio and the Zika epidemic, as can be observed from the number of bioRxiv submissions over time (see https://doi.org/10.1101/833400). Covid-19 does seem to be the event that drove their adoption in more clinically oriented fields as well, but it changes the rate of preprint submission much less drastically than the 2016 events.
3. More references could be cited in terms of previous articles comparing conclusions between preprints and published studies (e.g. Bero et al., 2021 - https://doi.org/10.1136/bmjopen-2021-051821)
Results:
Study design and LLM validation:
1. “The primary claims were classified into 6 types (mechanistic, associative, descriptive, methodological, therapeutic, or null result).” Please state upfront how was this done, and by whom (i.e. human, LLM, etc.) – it does become clear later but is unclear here. A definition of each category would also be appreciated, as well as the information of whether categories could coexist (e.g. it seems possible for a claim to be “mechanistic” and “null” at the same time, although from the GitHub prompt I got the impression that categories were mutually exclusive).
2. “We validated the model against four independent human raters” – on which tasks exactly? Claim parsing? Content change? Hedging shift? Claim classification? Please be more specific if the kappa and alpha values mentioned here relate to one or these categories, an aggregate of all of them or a particular subset of them. This should also be stated explicity in the legend for Figure S2b.
3. “Four-rater consensus” – how was consensus between raters reached? The Methods suggest this to be the mean ordinal rating, but “consensus” to me suggests that discrepancies were somehow resolved, so there may be a better way to describe this here (e.g. “mean rating of the four human raters”).
4. Although the agreement between the LLM and humans is generally reasonable, there is a non-negligible fraction of abstracts with major disagreements (e.g. 2% with “major” vs. “unchanged”). Did you examine these cases to see whether there is any particular pattern of error? Perhaps a few examples would be useful as supplementary material to make the discussion of model accuracy more concrete.
5. “The model was within one ordinal level of the consensus on 98.9% of pairs”. This is a somewhat misleading statistic, as for a 3-category scale it’s effectively impossible to be more distant than this if any of the ratings is in the middle of the scale. An arguably fairer denominator would be of all the ratings marked at the ends of the scale (e.g. “unchanged” or “major”), how many are within one ordinal level (which in Figures 2a and 2e seem to be at 98 and 97%).
6. An estimate of test-retest reliability of the model itself would be welcome here. The Methods state that “three Sonnet runs agreed at k=0.75”, but this seems to belong in the results (and it’s not clear to which part of the analysis that number refers to).
Primary scientific claims usually persist…
1. When assessing claim type transitions, are we sure that we are actually looking at the same claim? Or could we actually be looking at the replacement of a claim when a category changes? If the intention was to measure whether a claim actually changed category, wouldn’t it be important to exclude cases where it was entirely substituted from the denominator. Note that this would require that the data on claim substitution is available, which I’m not sure is the case. If it is not, this should be mentioned as a limitation.
Patterns of revision vary with claim type…
1. “The first secondary claim…”. Is there a natural order between secondary claims (a first and a second)? This was never mentioned up to now. If there is, wouldn’t it be clearer to use “primary”, “secondary” and “tertiary”?
2. The numbers cited in the text refer to averages between primary and secondary claims that are not explicitly displayed in Figure 2b, which is a bit confusing. It would probably be useful to either include the averaged results in the figure or those broken down by primary/secondary in the text to avoid confusion.
3. “Importantly, major revision declined across years of preprint posting”. Beyond discussing this further (as noted in the major issues), please specify the statistical test that generated the quoted p value.
4. Why use octiles of journals rather than individual points for journal impact? This seems rather arbitrary (e.g. one could choose multiple ways to aggregate journals, leading to the possibility of data dredging) and leads to loss of information, so I don’t see the point of performing this kind of aggregation. Having statistics using the full set of journals seems clearly preferrable here (although log transformation is still desirable due to the skewed distribution of journal impact).
5. For the negative association of preprinting with retractions, please provide at least an idea of how this analysis was performed in the results section – particularly how the comparator group of “papers that were never preprinted” was defined, as this is not clear at all.
6, A measure of Jaccard similarity appears in the supplementary analysis in Figure S5b. Why is this only used here? Wouldn’t it be interesting to show the overall distribution of Jaccard similarity somewhere in the main results (for instance, to allow the reader to see what proportion of abstract are identical or near-identical to the preprint version)?
Methods:
1. “The remainder failed on transient API limits”. Does this refer to Sonnet API limits? If so, why did the authors not simply run the model again in these abstracts?
2. “Up to two secondary claims” – once again, is this exactly two (as implied by the abstract) or up to two? Please standardize to whatever is correct.
3. “The consensus was defined by majority vote”. But what happens when there is no majority (which seems to be a possibility among four raters? Were these claims excluded from the analysis)? Note that this may bias the conclusion, as claims in which raters were divided are probably the harder cases where one would expect lower agreement of individual humans with LLMs as well.
4. “Three replicate Sonnet runs agreed at k=0.75” – again, for which task? Claim extraction, content change, hedging, or a combination of these? Please specify (although once again I think this belongs in the results).
Figures:
Figure 1b: the “2:1 cautious confident” in subpanel b is an interpretation and can probably be removed from a figure that already has a lot of information.
Figures 1 and S3. Claim category labels are not consistent (S3 has “Mechanistic”, “Methodological” and “Associative”, 1 has “Mechanism”, “Method” and “Association”. Please standardize for clarity.
Figure 2b: The fact that the X-axis scale only goes up to 15 does not allow the reader to visualize the proportion of change and leads one to overestimate differences. It seems better to have the scale go up to 100%, which would make the reader aware both of the infrequent proportion of changes and of the small difference between primary and secondary claims for most categories.
Still on this figure, also note that this type of plot suggests (at least to me) some kind of transition or continuity between the blue and red dots (likely because of the line that connects the to), but this is not the case (i.e. they refer to estimates from different samples), so I’d rather go without the line.
Figure 2e: once more, why use octiles of journals rather than individual points for journals? This is counterintuitive (to the point that it has to be explained in the legend) and leads to loss of information, so I don’t see the point of performing this kind of aggregation.
O autor declara que não possui conflitos de interesse.
The author declares that they did not use generative AI to come up with new ideas for their review.
No se han publicado comentarios aún.