**Background:** Gastroesophageal reflux disease (GERD) is a common condition with a global prevalence of 8–33%. While pH-impedance monitoring is the standard for confirming or excluding pathological GERD, interpretation of several key metrics—including the number of reflux episodes, post-reflux swallow-induced peristaltic wave (PSPW) index, and baseline impedance—requires time-consuming manual analysis. Artificial intelligence (AI), particularly machine learning and deep learning with convolutional neural networks (CNNs), has shown promise in automating the analysis of complex physiological signals. This review summarizes the current evidence on AI applications for pH-impedance monitoring in GERD diagnosis.
**Methods:** The authors conducted a narrative review of the available literature on AI applications in pH-impedance monitoring for GERD. They describe the general principles of AI, machine learning, and deep learning, and outline five key steps for developing clinically useful AI models: data collection and annotation, model development, interpretability, clinical validation, and collaboration with healthcare professionals. The review focuses on three specific AI applications: measuring the number of reflux episodes, measuring baseline impedance, and measuring PSPW index.
**Key Results:**
1. **Number of reflux episodes:** Two pilot studies were identified. The first study manually identified 2049 impedance events and used a decision tree algorithm with 24 nodes over nine layers, achieving 88.5% accuracy in identifying impedance events (reflux, air events, swallows, and artifacts). The second study utilized 7939 impedance events from 106 patients (aged 20–65 years with typical GERD symptoms and negative endoscopy) and employed a deep residual learning model with 18 convolutional layers (ResNet18). This model achieved 87% accuracy in recognizing reflux episodes, with an excellent inter-rater agreement (ICC = 0.965) between the AI model and manual annotation for calculating the number of reflux episodes per patient.
2. **Baseline impedance:** One study evaluated AI-augmented baseline impedance analysis using the same proprietary AI software. The AI extracted a novel parameter called artificial intelligence baseline impedance (AIBI) in both upright and supine positions by removing all artifactual interruptions (reflux, air, and swallowing events). A new metric, the upright:recumbent AIBI ratio (AIBI ratio), was investigated. The recumbent AIBI correlated with mean nocturnal baseline impedance (MNBI), but upright AIBI did not. Only the AIBI ratio differed between management responders and non-responders (defined as ≥50% symptom improvement), demonstrating a diagnostic advantage over other metrics. The AIBI ratio showed an area under the ROC curve of 0.661 for predicting response to medical therapy, compared to 0.715 for acid exposure time (AET).
3. **PSPW index:** A study involving 106 patients and 3761 reflux episodes used deep residual learning with 18 convolutional layers (ResNet18) to recognize PSPW events. Three experts manually annotated PSPW occurrence after reflux episodes for training and verification. The AI model achieved 82% accuracy in recognizing PSPW events, with an excellent inter-rater agreement (ICC = 0.921) between the AI model and manual annotation for calculating the PSPW index per patient.
**Clinical Implications:** The Lyon Consensus recommends AET <4% as normal and >6% as pathological, with values between 4% and 6% considered inconclusive. In these ambiguous cases, adjunctive metrics—including number of reflux episodes (<40 normal, >80 pathological), PSPW index (threshold of 61% to distinguish GERD from healthy controls), and MNBI—are needed. AI has the potential to automate the assessment of these metrics, which currently require time-consuming manual evaluation. The AIBI ratio may offer additional predictive value for treatment response. However, current limitations include limited training data, lack of compatibility validation across different pH-impedance analysis software, absence of clinical validation studies on diverse patient populations, lack of interpretability research, and accuracy below 90%. The authors emphasize that AI is not intended to change the diagnostic process for GERD but to provide complete physiological information more efficiently. Collaboration between computer scientists, engineers, and medical experts is deemed crucial for future development.