Health-Related Data Sources Accessible to Health Researchers From the US Government: Mapping Review
Journal of Medical Internet Research · 4 authors, 4 centres
AI SUMMARY
FIDELITY 100%
This summary was generated by AI from a single paper. It has not been reviewed by a clinician and is not clinical advice. Verify against the source before acting on it.
This systematic mapping review identified 57 federally sponsored, health-related US data sources accessible to researchers, with 75% offering free data sets. Most sources (53%) provided survey or assessment data, and 70% focused on individuals or patients, collecting demographic (77%) and clinical (61%) information. The review highlights a need for improved data standardization across government entities to facilitate researcher access and data integration.
Full summary
5,139 CHARS
**Background:** Big data from large, government-sponsored surveys and data sets offers researchers opportunities to conduct population-based studies of important health issues in the United States, as well as develop preliminary data to support proposed future work. Yet, navigating these national data sources is challenging. Despite the widespread availability of national data, there is little guidance for researchers on how to access and evaluate the use of these resources. The objective of this study was to identify and summarize a comprehensive list of federally sponsored, health- and health care–related data sources that are accessible in the public domain in order to facilitate their use by researchers.
**Methods:** The authors conducted a systematic mapping review of government sources of health-related data on US populations with active or recent (previous 10 years) data collection. A research librarian conducted an internet search of major federal agencies and subagencies involved in health research, including the Department of Health and Human Services, CDC, NIH, HRSA, FDA, CMS, SAMHSA, AHRQ, the Bureau of Labor Statistics, and the Census Bureau. The search was developed and conducted in 2019 and repeated in May 2021. Inclusion criteria required that data sources be federally sponsored by a US government agency, include data on US populations (individuals, patients, providers, or health care sites/systems), contain health- or health care-related data, and have active/ongoing data collection or data collection as recent as 2010 or later. Data sources were excluded if limited to biological/genetic samples, radiologic images, or laboratory tests; did not produce publicly available data sets; referred only to a parent entity; included only aggregated statistics; or were duplicates. Key measures abstracted included government sponsor, overview and purpose, population of interest, sampling design, sample size, data collection methodology, type and description of data, and cost. Convergent synthesis was used to aggregate findings. To ensure interrater reliability, team members reexamined at least 10% of all completed reviews, and any disagreements were discussed until 100% agreement was attained.
**Key Results:** Among 106 unique data sources identified, 57 met the inclusion criteria. Data sources were classified as survey or assessment data (n=30, 53%), trends data (n=27, 47%), summative processed data (n=27, 47%), primary registry data (n=17, 30%), and evaluative data (n=11, 19%). Most (n=39, 68%) served more than one purpose. The population of interest included individuals/patients (n=40, 70%), providers (n=15, 26%), and health care sites and systems (n=14, 25%). The sources collected data on demographic (n=44, 77%) and clinical information (n=35, 61%), health behaviors (n=24, 42%), provider or practice characteristics (n=22, 39%), health care costs (n=17, 30%), and laboratory tests (n=8, 14%). Most (n=43, 75%) offered free data sets. The majority (36/57, 63%) of data sources produced annual data sets, and most (49/57, 86%) were active with ongoing data collection. Sample sizes ranged from 637 (Compendium of US Health Systems) to over 100,000,000 (National Death Index), with approximately half having sample sizes larger than 50,000. The Department of Health and Human Services represented the largest share, covering 7 different entities and 52 data sources. The authors identified 8 families of related data sources. Three major findings emerged: (1) there is a vast amount of health-related big data available, accelerated by nationwide efforts to promote transparency and replicability; (2) it is easier to select a data source first and then ask a question, rather than asking a question first and finding a data source to answer it; and (3) there is a lack of structure and standardization of data across national data sources, limiting integration.
**Clinical Implications:** This review offers insight into the availability of large, federally funded sources of health data that can support a wide range of research efforts, including preliminary or pilot work, evidence for clinical decision-making, trends analyses, and outcome studies. National data sources offer several advantages over primary data collection, including standardized and cleaned data, large nationally representative samples enabling examination of rare events, and elimination of the financial burden of primary data collection. However, researchers should be aware of limitations including long lag times between data collection and release (often years), inability to follow distinct individuals longitudinally in many data sources, and the challenge of fitting predetermined research questions to pre-existing data. The authors note that further advancements are needed to catalog available data, unify data set presentation conventions, and facilitate retrieval of important aspects such as variable definitions. Despite these challenges, these resources provide academicians, clinicians, and researchers options at a lower cost and with more efficiency than primary prospective studies.