**Background:** Hypertension is a major risk factor for cardiovascular disease, COVID-19 complications, and cognitive impairment, affecting approximately 90% of individuals in industrialized countries over their lifetime. While genetic and environmental factors are known contributors, the underlying mechanisms remain poorly understood. Epigenetic changes, particularly DNA methylation, have been implicated in hypertension, but research in this area is less developed compared to other diseases such as cancer. This study aimed to determine whether clear epigenetic signatures detectable via DNA methylation and machine learning techniques could distinguish between healthy, pre-hypertensive, and hypertensive individuals.
**Methods:** The study used publicly available DNA methylation data from the GEO database (accession GSE 193795), comprising 132 individuals: 44 control (healthy), 44 pre-hypertensive, and 44 hypertensive patients, all aged 50–65 with equal sex distribution. Control subjects had systolic blood pressure ≤120 mmHg and diastolic ≤80 mmHg, normal blood lipids (cholesterol <5.18 mmol/L, triglycerides <1.70 mmol/L), and BMI 24.1 ± 2.5. Hypertensive patients had systolic ≥160 mmHg and diastolic ≥110 mmHg (or ≥140/90 mmHg on medication), BMI 27.3 ± 3.2. Pre-hypertensive patients had systolic 120–139 mmHg and diastolic 80–89 mmHg, BMI 25.3 ± 3.3. All patients met these criteria for 2 years before inclusion. For each individual, 223,945 CpG DNA methylation levels were obtained from peripheral blood using the Illumina protocol. A neural network with one hidden layer containing 50 artificial neurons was used for classification. Preliminary filtering selected the top 1% of CpGs (2,239) based on correlation with the classification variable. Secondary filtering further reduced CpGs using standard deviation, interquartile range, and range metrics. An optimization algorithm iteratively removed CpGs while maintaining accuracy. Training used 75% of data, testing used 25%, with 10-fold cross-validation and 100 simulations per configuration.
**Key Results:** The base model using 2,239 CpGs (top 1% by correlation) achieved a mean accuracy of 86.3% for distinguishing control from hypertensive/pre-hypertensive patients. Further filtering using standard deviation as a metric achieved a statistically comparable mean accuracy of 83.3% using only 22 CpGs. For distinguishing hypertensive from pre-hypertensive patients, the base model again generated accurate results, but secondary filtering metrics (interquartile range, range, standard deviation) resulted in statistically significant accuracy decreases. The optimization algorithm reduced the number of CpGs to 1,120 while achieving a mean accuracy of 88.3%, comparable to the base model's 91.9% accuracy for this more challenging classification task. The 22 CpGs identified in the reduced model are listed in Table 2 of the paper. These results represent a substantial improvement over prior work by Nguyen et al., who reported 69% accuracy for hypertension detection using DNA methylation and machine learning.
**Clinical Implications:** The study demonstrates that DNA methylation signatures from peripheral blood can serve as objective biomarkers for hypertension and pre-hypertension, unaffected by transient factors such as stress or recent exercise that can influence blood pressure measurements. This approach could be integrated into routine blood tests to screen for undiagnosed hypertension, potentially identifying the large number of individuals unaware of their condition. The ability to distinguish between pre-hypertensive and hypertensive states suggests potential for monitoring disease progression. The use of peripheral blood rather than cardiac tissue makes this a practical clinical tool. Future applications could include personalized medicine approaches where DNA methylation profiles guide treatment selection, as suggested by prior pharmacoepigenetic analyses. Limitations include the relatively small sample size (132 individuals), the cross-sectional design, and the 'black box' nature of neural network models, though the authors mitigated this through extensive CpG selection analysis. Larger longitudinal studies are needed to validate these findings and develop clinically applicable prediction models.