PREreview del Preregistration Works: Increased Reporting Quality, Internal Validity, and Protocol Adherence in Animal Studies
- Publicado
- DOI
- 10.5281/zenodo.21782911
- Licencia
- CC BY 4.0
PREreview of Preregistration Works: Increased Reporting Quality, Internal Validity, and Protocol Adherence in Animal Studies
Authored by: Sandra Grinschgl, Mwafaq Ramzi Haji, Haofu Huang, Fallon Mody, Chalermchai Rodsangiam, Max F. Wan, Quratul Ayn Zahara.
Summary:
The manuscript reports the results of a comparative study, examining the quality of reporting between studies with and without preregistration in experimental animal research, a field where the practice of preregistration has only recently gained traction. The study compares a pool of 22 preregistered published papers with a pool of 13 author-matched and 20 journal-matched publications, reporting that preregistered papers scored higher on reporting quality compared to the controls (mean=0.794 vs 0.637 author-matched (n=13) and 0.593 journal matched (n=20), p<0.001). This study adds to the growing body of literature evaluating the efficacy of preregistration as a metascience intervention.
Major comments:
1. The findings do not logically support the strong headline conclusions drawn, that “preregistration works”. The study demonstrates a correlation (preregistered papers → better reporting and internal validity) but cannot establish a direct causal link between preregistration and reporting quality due to a combination of small sample size and multiple confounding factors. The use of author- and journal-matched control groups do not fully account for potential confounds, including: self-selection bias (teams that voluntarily preregister may already have stronger open science culture), difference in topic complexity and methodological rigour of studies, shifts in field and journal reporting norms, whether compliance with basic ARRIVE guidelines is a good proxy for internal validity, and/or authors’ independent improvement of reporting standards over time.
The design supports that preregistered papers scored better than matched controls, but does not provide strong evidence that preregistration itself caused the better scores. The use of author- and journal-matched controls is a useful design feature, but we recommend the authors avoid implying that the two control groups cleanly isolate “content” versus “research team.” Journal-matched controls partly control for topic and publication venue, but they may also reflect journal-specific ARRIVE enforcement. Author-matched controls partly control for team culture and reporting habits, but not necessarily for study type, journal policy, or temporal changes in reporting practice. This limitation should be stated explicitly.
Therefore, we recommend the paper’s claims about preregistration be softened, and/or drawing out and discussing comparable findings on preregistration from other disciplines, as well as reflecting on possible improvements to animal research disciplinary norms to improve reporting quality and internal validity. We also recommend including effect sizes for the reporting of “effects of preregistration on reporting quality”: (The model revealed a significant main effect between the preregistered group and the controls (F(2, 37.957)=16.852, p<0.001), indicating that scores differed significantly across groups.)
2. A major but under-explored finding appears to be that preregistration does not consistently prevent undisclosed and/or unjustified deviations, notably that 21.5% of deviations were not disclosed and only 3.4% of deviations were justified. We suggest that the authors have the opportunity to expand on what kind of improvements they would like to see in preregistration templates. The study highlights that the two platforms ASR and PCT have different templates, and some statistical items are missing from one of the templates. Are there other improvements that the authors would like to see? Or are there issues in adoption of the templates? Is this about accessibility or training about how to use the templates?
3. Scope to improve discussion on ARRIVE essential 10 guidelines and scoring of reporting quality, and recommendations that follow from the findings. It is unclear if the authors developed the scoring system for reporting quality, we assume so. Similar studies in social sciences used more coders differences in the methods of those studies could be discussed. These studies include, for example,
Van den Akker, O., et al (2024). The potential of preregistration in psychology: Assessing preregistration producibility and preregistration-study consistency. Psychological Methods. https://doi.org/10.1037/met0000687
Hahn, L. et al (2025). A Cross-Sectional Study of the Completeness of Preregistrations by Psychological Authors From German-Speaking Institutions. Advances in Methods and Practices in Psychological Science, 8(3): 25152459251357568. https://doi.org/10.1177/25152459251357568
The authors could further elaborate that ARRIVE essential 10 are considered the basic minimum standards for research involving live animals. The ARRIVE guidelines website also provide a "Recommended Set" that are additional guidelines to reporting, encouraging researchers to report. Some topics in the Recommended Set are specific to animal research, such as "Housing and Husbandry", "Animal Care and Monitoring", while others are directly relevant to research quality and dissemination of results, e.g. "Interpretation / scientific implications" and "Generalisability". Would the authors encourage future researchers to report those as well? Are there subfield-specific guidelines to follow?
Additionally, would the authors recommend future meta-analysis to replicate this study using the same scoring for report quality, or should the scoring method be further refined and become more granular? Did the authors apply Regcheck to their Phase 2 and compare it to their ratings - how high is the agreement?
Minor comments:
1. The preprint reports that “Preregistered papers scored at the verge of the ‘Excellent’ category” in the Discussion, whereas it appears that the results place it in the upper part of the Good category?
2. The study’s preregistration was updated on scoring of reporting quality relevant to this section: The quality coefficient was categorised a priori as follow: ‘Excellent’ (0.8-1), ‘Good’ (0.6-0.79), ‘Average’ (0.4-0.59), ‘Mediocre’ (0.2- 0.39), and ‘Poor’ (0-0.19), as inspired by previous literature (26). It is unclear what was updated and if it was done before or after coding started.
3. Unclear if Phase 1 papers are also included in Phase 2 for the reporting in the section, “Sample description for the evaluation of the reporting consistency and protocol adherence (phase 2)”.
4. The Introduction could be improved with some additional background on the specific field of experimental animal research, for example, a sense for how many animals are involved in lab-based research worldwide or in the country.
Acknowledgements:
The authors did not exhaustively review the appendices.
Competing interests:
The authors declare they have no competing interests.
Competing interests
The authors declare that they have no competing interests.
Use of Artificial Intelligence (AI)
The authors declare that they did not use generative AI to come up with new ideas for their review.