Early Discrimination of Benign and Malignant Ovarian Tumors Using Explainable Machine Learning on Routine Blood Biomarkers
Eastern Journal of Medicine, cilt.31, sa.3, ss.519-526, 2026 (Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 31 Sayı: 3
- Basım Tarihi: 2026
- Doi Numarası: 10.5505/ejm.2026.36675
- Dergi Adı: Eastern Journal of Medicine
- Derginin Tarandığı İndeksler: Scopus, EMBASE, TR DİZİN (ULAKBİM), Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest), Pharma Collection (ProQuest)
- Sayfa Sayıları: ss.519-526
- Anahtar Kelimeler: biomarkers, calibration, decision curve analysis, explainable artificial intelligence, machine learning, ovarian cancer, SHAP
- Acıbadem Mehmet Ali Aydınlar Üniversitesi Adresli: Evet
Özet
Ovarian cancer remains one of the deadliest gynecological malignancies, largely due to delayed diagnosis and the limited specificity of individual tumor markers. This study aimed to develop an explainable machine learning (ML) framework for preoperative discrimination of benign and malignant ovarian tumors using routine laboratory data. The publicly available Soochow University dataset comprising 349 patients (171 benign, 178 malignant) with 49 clinical features was analyzed. To prevent data leakage, the cohort was first partitioned into tr aining and hold-out test sets, and all preprocessing was learned on the training partition only. Multi-stage missing-value handling-including multivariate imputation by chained equations (MICE) estimated exclusively from the training rows-yielded a final cohort of 345 patients and 48 features. Four classifiers-Random Forest, XGBoost, LightGBM, and Elastic Net-were trained using 5-fold repeated cross-validation and evaluated on an independent test set (n = 103). Discrimination was compared with the DeLong te st; model interpretability was assessed via SHAP analysis; and clinical utility was evaluated through numerical calibration metrics, reliability curves, and decision curve analysis (DCA). Random Forest achieved the best overall discrimination (AUC = 0.918, 95% CI: 0.860–0.976; sensitivity = 90.6%; specificity = 80.0% at a 0.50 cut-off) and was the best-calibrated model (calibration slope = 1.35, intercept = 0.08, Hosmer–Lemeshow p = 0.14). Pairwise DeLong tests indicated that Random Forest significantly outperformed XGBoost (p = 0.03), whereas its differences from LightGBM and Elastic Net were not statistically significant. SHAP analysis identified HE4, albumin, age, menopausal status, and CEA as the most influential predictors. DCA confirmed net clinical benefit across threshold probabilities ranging from 0.10 to 0.70. This explainable ML framework demonstrates that transparent models utilizing routine preoperative laboratory data can accurately and reliably differentiate benign from malignant ovarian tumors. The proposed approach ma y support clinical decision-making and preoperative risk stratification in patients with adnexal masses.