**Background:** DNA polymorphisms and self-defined ethnicity are used to group individuals by ancient geographic ancestry. Previous work has shown that ethnicity-related differences exist in newborn screening (NBS) metabolite levels for conditions like cystic fibrosis and congenital hypothyroidism. This study investigated whether ancestral relationships could be identified from routine newborn metabolic screening data, and whether metabolic differences between populations correlate with genetic and geographic distances.
**Methods:** This cross-sectional study analyzed 503,935 screen-negative singleton newborns from the California NBS program (2013–2017). After exclusions (unknown collection timing, extreme birth weight or gestational age, positive/unknown TPN status, multiple or unknown ethnicity), 392,303 infants remained across 17 parent-reported ethnic groups. Data included 41 blood metabolites measured by tandem mass spectrometry and 6 covariates. A 41×17 data matrix of mean metabolite levels was created and standardized. Manhattan distances between all 136 group pairs were computed. Genetic distances (Tau statistic) were calculated using 55 ancestry-informative SNPs for 15 of the 17 groups. Geographic distances were estimated using Google Maps with single geographic reference points per group. Hierarchical clustering, multidimensional scaling (MDS), and phylogenetic tree analysis (Neighbor-Joining method) were used to compare distance measures. Machine learning (logistic regression with L1 penalty/Lasso, 10-fold cross-validation, 20 repeats) was used to predict ethnic group association from individual metabolic profiles, with performance measured by AUC.
**Key Results:** Ethnicity-associated differences were found for 70.7% of NBS metabolites (29 of 41, Cohen's d > 0.2 in at least one group pair). The largest effect sizes were for C3 between Japanese and Samoan (d = 1.02), C2 between Chinese and Asian East Indian (d = 0.96), and C5OH between White and Black (d = 0.95). Acylcarnitine levels showed significantly higher variability between ethnic groups than amino acids (P < 1e-4). Hierarchical clustering identified three major metabolic clusters: (1) Japanese, Filipino, Korean, Chinese, Vietnamese; (2) Cambodian, Laos, Other Southeast Asian, Black; (3) Native American, Hawaiian, Middle Eastern, White, Guamanian, Hispanic. Samoan infants had higher levels for 33 of 41 metabolites (P < 0.001). Correlation between metabolic and genetic distances was low (Cor = 0.29), as was metabolic-geographic correlation (Cor = 0.27). Genetic-geographic correlation was higher (Cor = 0.54). Outlier analysis revealed Black-White pairs had larger genetic but smaller metabolic distances, while Chinese-Japanese pairs had smaller genetic but larger metabolic distances. Machine learning achieved highest AUC for Black vs. Chinese (0.96) and Korean vs. Native American (0.91), and lowest for Hispanic vs. Native American (0.51) and Cambodian vs. Laos (0.52). For genetically similar pairs (Tau ≤ 0.05) with larger metabolic differences (AUC ≥ 0.7), top informative metabolites included C10:1, C12:1, C3, C5OH, and leucine-isoleucine—all primary NBS markers for inborn metabolic disorders.
**Clinical Implications:** This study demonstrates that routine newborn screening data can reveal population-level metabolic differences that correlate—albeit weakly—with genetic ancestry. The finding that 71% of NBS metabolites vary by ethnicity, and that disease biomarkers (C10:1 for MCADD, C12:1 for VLCADD, C5OH for 3-MCC deficiency, C3 for methylmalonic acidemia, leucine-isoleucine for MSUD) are among the most informative for distinguishing genetically similar populations, suggests that ethnicity-related metabolite variability could contribute to false-positive rates in newborn screening. Incorporating metabolic ancestry information may help reduce confounding in genetic disease screening. The study also highlights the complex interplay of genetic and environmental factors (e.g., diet, admixture) in shaping newborn metabolism, particularly for populations like Samoan (highest metabolite levels) and Hawaiian (admixed ancestry). Limitations include reliance on 17 predefined ethnicity categories, unknown admixture, separate samples for metabolic vs. genetic data, and crude single-point geographic approximations.