It found no significant difference in model accuracy when trained on synthetic data, real data, or a combination, suggesting synthetic data can be used as a substitute for training algorithms.
PLOS ONE · 2 authors, 2 centres
This summary was generated by AI from a single paper. It has not been reviewed by a clinician and is not clinical advice. Verify against the source before acting on it.
It found no significant difference in model accuracy when trained on synthetic data, real data, or a combination, suggesting synthetic data can be used as a substitute for training algorithms.
The study used data from the National Diet and Nutrition Survey (NDNS) to evaluate if synthetic datasets could train machine learning models to predict mean arterial blood pressure. After feature selection, four variables (age, sex, weight, height) were used. Three model types (Bayesian linear regression, random forest, neural network) were trained on four different datasets: a real training set (n=2408), two synthetic sets (n=2408 and n=4816), and a combined real-synthetic set (n=4816). All models were tested on the same real test set (n=424). The synthetic datasets showed high fidelity to the real data. Limitations include that the NDNS data is not representative of the general UK population and the study did not assess generalizability or re-identification risk. The findings support the potential of synthetic data as a privacy-preserving tool for training predictive health algorithms.