**Background:** Diabetic retinopathy (DR) is a leading cause of vision loss, and screening can prevent blindness. In low- and middle-income countries (LMICs), barriers include limited skilled personnel and lack of access to eye care. Artificial intelligence (AI) systems for grading retinal images may help overcome these constraints. While several AI systems have been validated, evidence from real-world settings, especially in LMICs, is limited. This study evaluated the diagnostic accuracy of Medios AI software integrated into a smartphone-based retinal camera within the existing DR screening programme in Dominica, a Caribbean island nation with an estimated diabetes prevalence of 17.7%.
**Methods:** This prospective, cross-sectional clinical validation study enrolled consecutive patients with diabetes aged ≥18 years attending mobile DR screening clinics in four health districts in Dominica from 5 June to 3 July 2021. After pupillary dilation (tropicamide 0.5% and phenylephrine HCL 5%), a minimum of two images per eye (one optic disc-centered, one macula-centered) were taken using a hand-held smartphone-based Fundus on Phone camera (Remidio). The field grader (senior Dominican screener-grader) performed DR grading at the point of care and made referral decisions. AI grading (Medios DR AI software, NM App V2.0) was deferred to avoid influencing clinical decisions. The AI provides a binary output: 'signs of DR detected' or 'not detected', with a threshold of moderate non-proliferative DR (NPDR) or worse per the International Classification of Diabetic Retinopathy (ICDR). Referable DR (RDR) was defined as moderate NPDR or worse, diabetic macular edema (DME), or ungradable image in either eye. Vision-threatening DR (VTDR) was defined as proliferative DR and/or DME. Sensitivity, specificity, PPV, NPV, and AUC with 95% CIs were calculated comparing AI to the field grader as gold standard. Interobserver agreement between the field grader and remote graders from the English National Screening Programme was assessed using Cohen's kappa.
**Key Results:** A total of 587 participants were included (mean age 64 years, range 26–94; 72.6% women; 97.1% Black Caribbean). Mean duration of diabetes was 12 years (SD 8.8). The field grader classified 72 participants (12.2%) as having ungradable images in at least one eye, with 52 (8.8%) ungradable in both eyes. Interobserver agreement for detecting any DR between field and remote graders was κ=0.69 (good agreement). Prevalence of RDR by field grader was 45.4% (95% CI 41.5%–49.5%) including all participants, and 40.1% (95% CI 36.0%–44.3%) excluding ungradable participants. For all participants (n=587), the AI had a sensitivity of 77.5% (95% CI 72.0%–82.3%), specificity of 91.5% (95% CI 87.9%–94.3%), PPV 88.4% (95% CI 84.1%–91.7%), NPV 83.0% (95% CI 82.0%–87.9%), and AUC 0.84 for detecting RDR. Excluding 52 participants with both eyes ungradable (n=535), sensitivity was 80.4% (95% CI 75.5%–86.3%), specificity 91.5% (95% CI 87.9%–94.3%), PPV 86.6% (95% CI 81.7%–90.3%), NPV 87.9% (95% CI 84.6%–90.6%), and AUC 0.96. When compared to remote graders (excluding 65 ungradable participants), AI sensitivity and specificity were both 83.7% (95% CI 75.6%–90.4%), with AUC 0.86. For VTDR (excluding ungradable participants), AI sensitivity was 89.2% (95% CI 82.8%–95.2%); specificity could not be calculated due to the binary output. Among 12 VTDR cases missed by AI, 7 had non-DR macular pathology prompting referral; excluding these, VTDR sensitivity rose to 95.2% (95% CI 90.7%–99.3%).
**Clinical Implications:** This real-world study demonstrates that AI can be deployed in a mobile DR screening programme in an LMIC with reasonable accuracy compared to a trained specialist grader. The sensitivity (77.5%–80.4%) was lower than in prior controlled studies (93%–100%), likely due to real-world conditions including variable image quality and the field grader's ability to detect non-DR pathology. The high specificity (91.5%) suggests appropriate referral decisions, reducing unnecessary burden on eye clinics. The AI's performance for VTDR (89.2% sensitivity) was promising, especially after accounting for non-DR referrals. The study highlights the importance of integrating AI with existing workflows and the need for adequate training and quality assurance. Limitations include the single-country setting, predominance of women (possibly due to differential healthcare access), and the inability to use AI's real-time image quality feedback during the study. Future research should evaluate AI operated by community nurses to further expand screening coverage.