Commentaires
Écrire un commentaireAucun commentaire n’a encore été publié.
This is an evidence map of 11,628 PubMed-indexed records on generative language models in healthcare, January 2023 to June 2026, built to answer one question: is the clinical literature keeping pace with the systems it evaluates? The central construct is evaluation lag, the number of quarters between the release of the newest model family a study names and that study's own publication quarter. Mean lag widens from 1.33 to 6.08 quarters. Because a discontinued family ages at exactly one quarter per quarter, the authors benchmark that rise against a counterfactual holding 2023 model composition fixed, and find that migration to newer systems offsets 56.2% of the drift (95% CI 49.9-65.2, 2,000 bootstrap resamples). By design, randomised controlled trials evaluate models 4.62 quarters older than other empirical designs (95% CI 3.62-5.63), yet among the records naming a family still receiving releases no design differs from any other; what separates designs is the probability of studying a discontinued family at all, 62.1% for randomised trials against 15.6% for preprints.
The framing is right and the execution is unusually careful for a bibliometric paper. Reporting the mechanical ageing rate as the null instead of zero, validating the decomposition by showing that discontinued-only records widen at 1.062 quarters per quarter, and running a pessimistic sensitivity analysis that assumes the earliest release of every named family are all things this literature normally omits. The version-specification result in Section 2 deserves to be read on its own: randomised trials specify the model version in 50.0% of cases, the lowest of any empirical design, which means half of the highest-tier evidence cannot be attributed to a determinate system. My comments below are about three places where the reported precision outruns the design.
1. Evaluation lag is measured on a differentially selected half of the corpus, and the headline comparison is between designs that differ in exactly that selection. Lag is computable for 5,775 of 11,628 records, and coverage by design runs from 87.0% for comparative evaluations and 69.0% for randomised trials down to 36.2% for preprints and 27.0% for reviews. The paper states this plainly, which is to its credit, but the between-design contrast is its central result and the missingness is both large and design-dependent. The direction is not obvious either: if preprints name a model in the abstract mainly when that model is the novelty on sale, the preprint advantage is inflated; if trials omit the name when the system is an incidental component, the trial penalty may be understated. Two things would settle it. Worst-case bounds in the Manski sense across the plausible range would show whether the ordering survives at all, and a full-text audit of a random sample, a few hundred records stratified by design, would give an empirical estimate of how the non-naming records differ. Without one of those, the contrast rests on an assumption of ignorable missingness that the paper does not state.
2. The abstract quotes the specification that maximises the effect and omits the one that does not survive correction. The median regression gives 4.62 quarters at P = 3.4 x 10^-19; ordinary least squares on the same specification gives an adjusted difference of 1.57 quarters, P = 0.021, and after Holm correction across the seven design contrasts P = 0.084. Median regression is the defensible choice for a strongly right-skewed distribution and the paper says so, so this is not a case of specification shopping. It is a case of a reader meeting one number and not the other. The magnitude differs by a factor of about three between estimators, and the abstract should carry that range, or at minimum state that the estimate is estimator-sensitive while the ordering is not.
3. The mechanism claim is carried by 22 records, and the abstract states it as a finding rather than as the absence of evidence. Among studies naming an actively developed family, randomised trials differ from reference by 0.02 quarters (95% CI -0.50 to 0.55, n = 22). Section 2 is careful here, noting that this establishes that the gradient operates through model selection rather than demonstrating the absence of a residual timeline effect. The abstract is not: "among studies naming a model still under development no design differed from any other" reads as an equivalence result. With n = 22 the interval cannot exclude differences of half a quarter in either direction, which is not nothing when the effect being explained is a few quarters. State the equivalence bounds the sample supports, and match the abstract to the caveat already in the body.
4. The counterfactual is anchored on a very thin quarter. The 56.2% offset depends on holding model composition at its 2023 value, and 2023-Q1 contains 52 records with five distinct families, a Herfindahl-Hirschman index of 5,424 and 72.0% of mentions on one family. The paper already shows the sensitivity by reporting 64.7% when the first quarter alone is the baseline, which is a nine-point swing from the headline. Report the offset as a function of the baseline window, at least 2023-Q1, 2023-H1 and 2023 in full, each with its bootstrap interval, so the reader can see how much of the 56% is a property of the literature and how much is a property of one small quarter.
The paper claims that harvesting, screening, classification and analysis execute end to end from deposited code, and that is the right standard for an evidence map that is meant to be regenerated as the literature grows. I could not find the archive identifier in the text I read, so the first thing to add is a repository DOI with a commit hash, alongside the verbatim PubMed query strings for all fifteen query blocks and the harvest date. The literature moves; a map without its query date cannot be reproduced even by its own authors.
Two derived artefacts deserve to be released as data in their own right. The first is the model-family release-date table that defines the lag variable: every number in the paper depends on it, and it would be reusable by anyone studying model currency in any field. The second is the regular expression set used for model detection, together with the precision audit. The audit reports that 9.9% of matches were flagged for the eleven most exposed families, or 2.0% of the corpus, which is a precision estimate. Recall is not estimated anywhere, and it matters more here, because a study that names its model only in the methods section is invisible to an abstract-level regular expression, and whether that happens is plausibly design-dependent. A recall estimate from the same full-text sample proposed in point 1 would cover both issues at once.
Add worst-case bounds or a full-text audit for the 50.3% of records without a computable lag, reporting coverage-adjusted design contrasts; state the missingness assumption explicitly.
Put both the median-regression and the Holm-corrected OLS estimates for the randomised-trial contrast in the abstract, or state that the magnitude is estimator-sensitive.
Report equivalence bounds for the active-family subgroup and align the abstract with the n = 22 caveat already given in the body.
Report the counterfactual offset across several baseline windows with intervals, rather than one figure from a 52-record quarter.
Give the repository DOI and commit hash, the fifteen query strings verbatim, and the harvest date.
Release the model-family release-date table and the detection patterns; add a recall estimate for model detection, reported by design.
State how the 12.0% of records with unattributable first affiliation are handled in the regional comparison, since the trial-grade shares compared there are 1.5% to 3.9% and that missing share could move them.
The finding that computer science venues and preprint servers contain no trial-grade record should carry the PubMed-coverage caveat in the same sentence, not only in the limitations paragraph, since it is the kind of line that travels alone.
Consider replacing the 45-fold growth headline with the incidence rate ratio of 1.272 per quarter (95% CI 1.257-1.289). The fold change compares two single quarters and the base quarter has 52 records; the IRR is the estimate that carries an interval.
A P value of 3 x 10^-19 conveys nothing beyond the effect and its interval; consider reporting the estimate with its CI and dropping the exponent.
This should be published after revision. The central quantity is well chosen, the counterfactual is the right way to handle mechanical ageing, and the version-specification result is a concrete, actionable finding that the reporting guidelines the paper cites could enforce tomorrow. What needs work is the treatment of the half of the corpus where the key variable is undefined, and the alignment of the abstract with the uncertainty the body already acknowledges. The conclusion I would defend on this evidence is slightly narrower than the one stated, and no less interesting: the highest tier of clinical evidence is systematically about superseded systems, and the gap is driven by which systems trials select rather than by how long trials take.
Competing interests: none.
Evgenii Arsentev, PhD
The author declares that they have no competing interests.
The author declares that they used generative AI to come up with new ideas for their review.
Aucun commentaire n’a encore été publié.