» Articles » PMID: 37433624

Prevalence and Predictors of Data and Code Sharing in the Medical and Health Sciences: Systematic Review with Meta-analysis of Individual Participant Data

Overview
Journal BMJ
Specialty General Medicine
Date 2023 Jul 11
PMID 37433624
Authors
Affiliations
Soon will be listed here.
Abstract

Objectives: To synthesise research investigating data and code sharing in medicine and health to establish an accurate representation of the prevalence of sharing, how this frequency has changed over time, and what factors influence availability.

Design: Systematic review with meta-analysis of individual participant data.

Data Sources: Ovid Medline, Ovid Embase, and the preprint servers medRxiv, bioRxiv, and MetaArXiv were searched from inception to 1 July 2021. Forward citation searches were also performed on 30 August 2022.

Review Methods: Meta-research studies that investigated data or code sharing across a sample of scientific articles presenting original medical and health research were identified. Two authors screened records, assessed the risk of bias, and extracted summary data from study reports when individual participant data could not be retrieved. Key outcomes of interest were the prevalence of statements that declared that data or code were publicly or privately available (declared availability) and the success rates of retrieving these products (actual availability). The associations between data and code availability and several factors (eg, journal policy, type of data, trial design, and human participants) were also examined. A two stage approach to meta-analysis of individual participant data was performed, with proportions and risk ratios pooled with the Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis.

Results: The review included 105 meta-research studies examining 2 121 580 articles across 31 specialties. Eligible studies examined a median of 195 primary articles (interquartile range 113-475), with a median publication year of 2015 (interquartile range 2012-2018). Only eight studies (8%) were classified as having a low risk of bias. Meta-analyses showed a prevalence of declared and actual public data availability of 8% (95% confidence interval 5% to 11%) and 2% (1% to 3%), respectively, between 2016 and 2021. For public code sharing, both the prevalence of declared and actual availability were estimated to be <0.5% since 2016. Meta-regressions indicated that only declared public data sharing prevalence estimates have increased over time. Compliance with mandatory data sharing policies ranged from 0% to 100% across journals and varied by type of data. In contrast, success in privately obtaining data and code from authors historically ranged between 0% and 37% and 0% and 23%, respectively.

Conclusions: The review found that public code sharing was persistently low across medical research. Declarations of data sharing were also low, increasing over time, but did not always correspond to actual sharing of data. The effectiveness of mandatory data sharing policies varied substantially by journal and type of data, a finding that might be informative for policy makers when designing policies and allocating resources to audit compliance.

Systematic Review Registration: Open Science Framework doi:10.17605/OSF.IO/7SX8U.

Citing Articles

Systematic Review: AI Applications in Liver Imaging with a Focus on Segmentation and Detection.

Pomohaci M, Grasu M, Baicoianu-Nitescu A, Enache R, Lupescu I Life (Basel). 2025; 15(2).

PMID: 40003667 PMC: 11856300. DOI: 10.3390/life15020258.


Evaluating the Reproducibility and Verifiability of Nutrition Research: A Case Study of Studies Assessing the Relationship Between Potatoes and Colorectal Cancer.

Jamshidi-Naeini Y, Vorland C, Kapoor P, Ortyl B, Mineo J, Still L medRxiv. 2024; .

PMID: 39677420 PMC: 11643200. DOI: 10.1101/2024.12.01.24318272.


Ask, and it shall be given you - individual patient data and code availability for randomised controlled trials submitted for publication.

Bramley P Anaesthesia. 2024; 80(2):205-206.

PMID: 39638367 PMC: 11726263. DOI: 10.1111/anae.16503.


Indicators of transparency and data sharing in scientific writing in published randomized controlled trials in orthodontic journals between 2019 and 2023: an empirical study.

Schueller S, Mikelis F, Eliades T, Koletsi D Eur J Orthod. 2024; 46(6).

PMID: 39569723 PMC: 11579657. DOI: 10.1093/ejo/cjae064.


COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study.

Sofi-Mahmudi A, Raittio E, Khazaei Y, Ashraf J, Schwendicke F, Uribe S PLoS One. 2024; 19(11):e0313991.

PMID: 39556553 PMC: 11573139. DOI: 10.1371/journal.pone.0313991.


References
1.
Hamilton D, Fraser H, Hoekstra R, Fidler F . Journal policies and editors' opinions on peer review. Elife. 2020; 9. PMC: 7717900. DOI: 10.7554/eLife.62529. View

2.
Ascha M, Katabi L, Stevens E, Gatherwright J, Vassar M . Reproducible Research Practices in the Plastic Surgery Literature. Plast Reconstr Surg. 2022; 149(4):810e-823e. DOI: 10.1097/PRS.0000000000008956. View

3.
Haddaway N, Grainger M, Gray C . Citationchaser: A tool for transparent and efficient forward and backward citation chasing in systematic searching. Res Synth Methods. 2022; 13(4):533-545. DOI: 10.1002/jrsm.1563. View

4.
Naudet F, Sakarovitch C, Janiaud P, Cristea I, Fanelli D, Moher D . Data sharing and reanalysis of randomized controlled trials in leading biomedical journals with a full data sharing policy: survey of studies published in and . BMJ. 2018; 360:k400. PMC: 5809812. DOI: 10.1136/bmj.k400. View

5.
Womack R . Research Data in Core Journals in Biology, Chemistry, Mathematics, and Physics. PLoS One. 2015; 10(12):e0143460. PMC: 4670119. DOI: 10.1371/journal.pone.0143460. View