PREreview del Human-AI Collaboration for Estimating Scientific Replicability
- Publicado
- DOI
- 10.5281/zenodo.21482056
- Licencia
- CC0 1.0
Summary
Questions:
(RQ1) To what extent can hybrid human-AI prediction markets more accurately forecast the outcomes of scientific replications compared to artificial-only (agent) or human-only markets?
(RQ2) Which information and strategies do human participants use when trading in a replication prediction market?
Background: Large-scale replication projects across psychology, economics, sociology, and other fields have produced disappointing replication rates, motivating methods that prioritize which studies to replicate. Two dominant paradigms exist: human crowd forecasting (surveys, prediction markets, structured elicitation) and machine learning models trained on paper metadata, statistics, and text. Each has complementary weaknesses: human forecasts are subject to cognitive biases and limited exposure to the literature and automated models miss contextual credibility signals.
Methods: Agents were trained on 402 replication outcomes from major replication projects (RPP, SSRP, EERP, Many Labs 1 and 2) plus DARPA SCORE studies; 30 additional SCORE outcomes (unpublished at experiment time) were held out as test data, five per discipline across six domains.
Forty-one statistical, bibliometric, author, venue, and semantic features were extracted per claim. Six live 12-hour online market events were run in April 2023 with 97 researcher participants recruited from relevant disciplines. Agents received $500, whereas humans received $25 per market, a three-trade minimum for payout eligibility, and a randomly selected "money market" for incentive-compatible compensation. Economics, sociology, and psychology events included human-only, hybrid, and artificial markets; marketing, political science, and education included only hybrid and artificial markets. Post-experiment surveys probed trading strategies.
Results:
Hybrid markets matched or outperformed artificial markets in most domains, with marketing and education as exceptions.
Hybrid markets achieved the lowest mean absolute error in sociology (0.424) and political science (0.386).
Human-only markets performed best in psychology (0.378) and economics (0.414, though the hybrid's 0.411 was marginally lower).
Notably, the artificial market's final prices clustered in a narrow band (~0.62–0.76) and predicted "will replicate" for all 30 test claims. Surveys indicated participants traded primarily on epistemic beliefs about replicability rather than profit-maximizing strategies, with some trend-following and limited strategic engagement.
Implications: The authors argue hybrid markets offer a scalable framework for combining algorithmic pattern recognition with human contextual judgment in scientific evaluation, with potential applications in confidence assessment, replication prioritization, and research funding decisions.
Major Issues (by type)
Ethics/Disclosure (Important): No conflict-of-interest statement and no funding acknowledgment anywhere in the preprint (pp. 1–12), despite structural dependence on DARPA SCORE data. 89 of 402 training outcomes and all 30 test outcomes come from SCORE (p. 6, §3, paras. 1–2) and continuity with a SCORE-era predecessor system ([50]). If any of this work was DARPA-funded, that must be disclosed; if not, an explicit statement removes the ambiguity. Participant compensation ($40 × 97 plus money-market payouts, p. 9) likely had an external funder that readers should be able to identify. IRB approval is asserted without a protocol number or named institution (p. 9, §4.3, para. 2).
Rigor/Skepticism (Critical): The conclusion's claim of consistently matching or outperforming both baselines (p. 11, §6, para. 1) is contradicted by Table 4 (p. 11): human-only markets beat hybrids decisively in psychology (0.378 vs. 0.523), and artificial markets beat hybrids in marketing (0.430 vs. 0.490) and education (0.458 vs. 0.488). The word "consistently" and the phrase "human only baselines" should both be removed or heavily qualified. The abstract's more careful "except for a few cases" (p. 1) shows the authors know the result is mixed, so this is a fixable framing inconsistency between abstract and conclusion rather than a disagreement with the data.
Rigor (Critical): The artificial market predicts "R" (will replicate) for all 30 test claims (Tables 2–3, pp. 10–11), with final prices confined to a narrow ~0.62–0.76 band. A baseline that never predicts "not replicate" is behaviorally indistinguishable from a constant classifier at the base rate; outperforming it is weak evidence for hybrid value. This degenerate behavior is never reported, analyzed, or acknowledged, yet it underwrites every hybrid-vs-AI and human-vs-AI comparison in the paper. The authors should report the artificial market's confusion matrix, investigate whether it is a calibration or implementation artifact, and re-weight the comparative claims accordingly.
Rigor (Critical): No uncertainty estimates or statistical tests accompany any MAE comparison, all of which rest on five claims per domain (Table 4, p. 11). Differences as small as 0.411 vs. 0.414 (hybrid vs. human-only, economics) are presented as meaningful orderings. With n=5, none of the domain-level rankings can be distinguished from noise; at minimum the authors should report confidence intervals or bootstrap distributions, and preferably a pooled analysis across all 30 claims with a pre-specified primary metric (MAE or Brier score).
Rigor/Reproducibility (Critical): The results in Tables 2–4 cannot be computationally reproduced from what is deposited. There is no code for the platform, agents, or analysis; final hyperparameter values (lambda, liquidity, percent difference) are withheld despite the tuning procedure being described (p. 8, §4.1); no random seeds are given for the stochastic genetic-algorithm training; and no per-market transaction logs are shared. Only the inputs (41 features, train/test splits at OSF) are reproducible.
Rigor/Replicability (Critical/Important): The findings are not currently crisp enough to be replicable as claims: with no pre-specified primary outcome and no uncertainty estimates, a replication could not tell success from failure. Direct replication is further blocked because the SCORE outcomes are now public (p. 6, §3, para. 2), contaminating any rerun, and because the human-side protocol (recruitment materials, instructions, pre-study module, survey instrument) is unavailable. The authors' own §4.1 shows the market is hyperparameter-sensitive to the point of degenerate no-trade or all-trade regimes, so without the selected configuration a replicator cannot separate "hybrid markets work" from "this configuration worked once."
Rigor (Important): Several analytic definitions needed to interpret or reanalyze the results are unstated: the decision threshold converting continuous prices to R/NR predictions, the precise target of the MAE (binary outcome vs. replication effect size), and the base rate of the 30 test outcomes (which sets the score a trivial classifier would achieve). Without the base rate, readers cannot judge whether any market beats "predict the majority class."
Rigor (Important): Training/test domain shift is unexamined. The training corpus is ~63% psychology and ~25% economics (Table 1, p. 6), yet test domains include political science (6 training examples), education (5), and marketing (20). The artificial market's weakness in exactly the low-training-data domains is predictable and should be discussed; relatedly, the AI market posted its worst MAE (0.528) in psychology, its best-represented training domain, which deserves explanation rather than silence.
Rigor/Design (Important): The minimum-activity rule requiring three trades for payout eligibility (p. 8, §4) plausibly introduces a selection effect toward more engaged, more confident traders, and its influence on aggregated human judgment is neither analyzed nor acknowledged. Because the paper simultaneously reports that a non-trivial number of participants made no transactions (p. 11, §6, para. 2), the interaction between the activity rule, engagement heterogeneity, and market accuracy is a live confound the authors should address.
Transparency (Important): Beyond the reproducibility gaps, the paper omits the number of agents instantiated per market, the number of survey respondents, the survey instrument itself, and the qualitative coding procedure behind the RQ2 strategy categories (p. 11, §5.2). The survey findings are the entire evidentiary basis for RQ2, so their provenance and coding should be documented and the instrument shared.
Clarity (Important): Terminology drifts between "replicability" and "reproducibility" despite the definitional footnote that distinguishes them (p. 3, footnote 2). Section 4 describes agents that "predict reproducibility" (p. 7), Figure 3 outputs a "Reproducibility Score" (p. 8), and the platform UI labels assets "Reproducible/Non-Reproducible" (Figure 4, p. 9) while the text says contracts represent "will replicate"/"will not replicate" (p. 7, §4, para. 2). Given the care taken in the footnote, the inconsistency actively confuses the paper's central construct.
Clarity/Rigor (Important): The relationship between the 97 total participants (p. 9, §4.3, para. 2) and the per-event counts, which sum to 214 (40 + 33 + 40 + 33 + 37 + 31, pp. 9–10), is never explained. Substantial cross-event participation is implied but undescribed, which bears directly on the independence of the six events and on the meaning of "corresponding disciplinary backgrounds" (p. 2). Did participants trade outside their home discipline, and if so, how does that affect the "domain expertise" interpretation of the human contribution?
Transparency/Data (Important): Table 2 contains an internal inconsistency (p. 10): market x0pA (sociology) lists a human final price of 0.23 but a human final prediction of "R." One of the two cells is wrong, and because the value feeds the human-only MAE, the error is not cosmetic.
Applicability/Rigor (Important): The paper does not benchmark its MAEs against the accuracy figures reported in the cited prior human-market literature (e.g., the Camerer/Dreber replication-market projects) or against the authors' own supervised models trained on the same 41 features (Wu et al., [62]). Readers therefore cannot situate the reported performance relative to what either prior paradigm already achieved, weakening the paper's implicit claim to advance the state of the art.
Applicability/Skepticism (Important): The scalability claim in the conclusion (p. 11) is not costed. Recruiting 97 disciplinary experts, compensating them, and running six 12-hour live events is precisely the non-scalable component of the pipeline, yet the paper offers no cost-per-forecast comparison against purely automated alternatives. Because human recruitment is the bottleneck, "scalable framework" needs either supporting cost analysis or softening.
Minor Issues
Abstract (p. 1): "more accurate and reliable" — no reliability or calibration analysis (e.g., calibration curves, Brier decomposition) is presented anywhere, so "reliable" is unsupported; either add the analysis or drop the word.
Abstract vs. body: the abstract says hybrid markets "match or outperform artificial prediction markets" (p. 1) while §6 escalates to also outperforming human-only baselines (p. 11); align the two.
§6, para. 1 (p. 11): "This works puts forward" → "This work puts forward."
§5.1, Economics para. (p. 9): "as shown in in Table 4" — duplicated "in."
Tables 2–3 captions (pp. 10–11): "Final value (in dollars)" conflicts with Figure 2's note (p. 7) that prices are "multiplied by 100 and converted to dollars"; the tabled values are unit-interval prices/probabilities in [0,1], not dollars. Reconcile the caption, the figure note, and the table units.
Figure 1 caption (p. 6) ends with a stray period on its own line.
Figure 2 (p. 7): the notional feature-space illustration is not referenced by axis units or explained in enough detail for a reader to connect "projected feature 1/2" to the 41-feature schema in Figure 1; a one-line bridge would help.
Figure 3 (p. 8): the five-step schematic labels the output a "Reproducibility Score" (see Clarity major issue) and shows "86 — High Confidence"; clarify whether 86 is a price (¢), a probability, or a separate confidence metric, since none is defined in text.
Figure 4 (p. 9): the screenshot shows "Markets will close in about 14 minutes" and 7-second x-axes despite live markets running 12 hours; label it as a demo/beta capture to avoid implying markets were minutes long.
Recruitment window given as "between January and March 2023" (p. 9, §4.3, para. 2) but the earliest market (marketing) ran 12 April 2023 (p. 10); clarify whether recruitment continued past March or the window is approximate.
Market date ordering: results are narrated economics-first (p. 9) but the events ran marketing (12 Apr) → political science (14 Apr) → education (17 Apr) → economics (19 Apr) → psychology (21 Apr) → sociology (24 Apr); a sentence noting the narration order differs from the run order would prevent confusion.
"97 researchers actively participated" (p. 9) vs. the money-market profit counts reported per event (e.g., "20 out of 40," "14 out of 33," pp. 9–10); state once whether these denominators are per-event participants or per-market traders.
Reference list defects: [17] (Cowgill et al.) venue rendered as "In amma, page 3" appears garbled (p. 12); [27] (Gillen et al.) "In Working paper" is incomplete (p. 13); [22] (Fahse & Schmitt) lacks venue information (p. 13); several arXiv entries mix "arXiv preprint arXiv:NNNN" formatting inconsistently.
Inconsistent quotation marks around "will replicate"/"will not replicate" (straight vs. directional) throughout §4 (p. 7).
Inconsistent capitalization of "Political Science" vs. "Political SC" between the prose (p. 10) and Table 4 (p. 11).
Non-breaking spacing/typo artifacts from the HTML-to-text conversion appear in several author names in the reference list (e.g., diacritics rendered as "Bahn ˇ ´ık"); worth a pass before submission.
The paper alternates between "agents," "artificial agents," "algorithmic agents," "synthetic agents," and "bots" (§§1–4); standardizing one term would improve readability.
§4.3 (p. 9) states single-share trades for both agents and humans, but §4 (p. 7) describes agents purchasing "contracts" without noting the single-share restriction; consolidate the trading-mechanics description in one place.
Note: This Human-AI review is the result of ASAPBio PreReview (Meta-research club), a crowdsourced peer review initiative and members used models like Gemini, Claude Fable 5, Sonnet 5, and Opus 4.8 to review the preprint and summarize the comments.
Competing interests
The authors declare that they have no competing interests.
Use of Artificial Intelligence (AI)
The authors declare that they used generative AI to come up with new ideas for their review.