Korean J Intern Med > Volume 41(5); 2026 > Article
ORIGINAL ARTICLE
Korean J Intern Med. 2026;41(5):862-872.         doi: https://doi.org/10.3904/kjim.2024.245
Clinical feasibility of deep learning-assisted classification of Helicobacter pylori infection in endoscopic imagery: a reader study
Jun-young Seo1,2, Jiseon Kang3, Do Hoon Kim1 , and Namkug Kim4
1Department of Gastroenterology, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
2Digestive Disease Center, CHA Bundang Medical Center, CHA University School of Medicine, Seongnam, Korea
3Department of Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
4Department of Convergence Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
Corresponding Author: Do Hoon Kim  , Tel: +82-2-3010-3197, Fax: +82-2-3010-6517, Email: dohoon.md@gmail.com
Namkug Kim  , Tel: +82-2-3010-6573, Fax: +82-2-476-4719, Email: namkugkim@gmail.com
Received: July 14, 2024;   Revised: January 20, 2026;   Accepted: June 16, 2026.
Share :  
Abstract
Background/Aims: Detecting Helicobacter pylori infection through endoscopic imaging is a preferred method to address the limitations of invasive diagnostics. We have developed a deep learning (DL) model to classify these images based on their H. pylori infection status and to evaluate its potential as a diagnostic aid.
Methods: We retrospectively enrolled H. pylori-positive patients (by rapid urease test and serum IgG) at Asan Medical Center between January 2015 and December 2020, establishing development and test datasets to evaluate model utility. We developed a DL model and assessed its performance at both the image and patient levels using metrics such as accuracy, sensitivity, specificity, and F1-score. Fourteen readers of varying experience levels interpreted the endoscopic images in two apsessions, before and after the integration of the DL model results. The diagnostic accuracy of the readers was analyzed using the McNemar test and generalized estimating equations.
Results: In the utility test set, the image-level accuracy, sensitivity, specificity, and F1-score of our DL model were 84.2%, 74.4%, 93.9%, and 82.5%, respectively. At the patient level, the corresponding values were 97.0%, 100%, 94.0%, and 97.1%. Endoscopists utilizing the DL model demonstrated significantly improved accuracy in classifying H. pylori infections, with classification rates of 83.2% versus 72.4% (p < 0.001).
Conclusions: Our DL model shows exceptional predictive capabilities for identifying H. pylori infections in gastroscopy images, suggesting the potential of DL-based tools to enhance clinical decision-making in endoscopy. Future research should focus on multicenter prospective studies to validate these findings.
Keywords: Artificial intelligence ; Deep learning ; Endoscope ; Helicobacter pylori

Editorial Office
101-2501, Lotte Castle President, 109 Mapo-daero, Mapo-gu, Seoul 04146, Korea
Tel: +82-2-2271-6792    Fax: +82-2-790-0993    E-mail: kaim@kams.or.kr                

Copyright © 2026 by Korean Association of Internal Medicine.

Close layer
prev next