Hastalık önleme stratejilerinin 5 yaş altı ölüm oranını belirlemedeki rolü: Makine öğrenimi yaklaşımlarıyla karşılaştırmalı bir analiz
The role of disease prevention strategies in predicting under-five mortality rates: A comparative analysis of machine learning approaches
- Tez No: 1024685
- Danışmanlar: PROF. DR. EMRAH ÖNDER
- Tez Türü: Doktora
- Konular: Sağlık Yönetimi, İstatistik, Yönetim Bilişim Sistemleri, Healthcare Management, Statistics, Management Information Systems
- Anahtar Kelimeler: Belirtilmemiş.
- Yıl: 2026
- Dil: Türkçe
- Üniversite: İstanbul Üniversitesi
- Enstitü: Sosyal Bilimler Enstitüsü
- Ana Bilim Dalı: Sayısal Yöntemler Ana Bilim Dalı
- Bilim Dalı: Belirtilmemiş.
- Sayfa Sayısı: Belirtilmemiş.
Özet
Beş yaş altı ölüm oranı, bir ülkenin sağlık sisteminin etkinliğini, sosyoekonomik gelişmişlik düzeyini ve çevresel sağlık koşullarını yansıtan en önemli halk sağlığı göstergelerinden biridir. Dünya Sağlık Örgütü ve Birleşmiş Milletler tarafından belirlenen Sürdürülebilir Kalkınma Hedefleri kapsamında, 2030 yılına kadar beş yaş altı ölüm oranlarının önemli ölçüde azaltılması hedeflenmektedir. Bu hedefe ulaşılabilmesi için çocuk ölümlerini etkileyen sağlık, çevresel, demografik ve sosyoekonomik faktörlerin bütüncül bir yaklaşımla incelenmesi gerekmektedir. Bu çalışmanın temel amacı, ülkelerin hastalık önleme kapasitesini yansıtan çeşitli göstergelerin beş yaş altı ölüm oranı üzerindeki etkilerini makine öğrenmesi yöntemleri kullanarak analiz etmek, farklı modelleme stratejilerinin performanslarını karşılaştırmak ve elde edilen sonuçları model agnostik analiz teknikleriyle yorumlamaktır. Çalışmada Dünya Bankası veri tabanından elde edilen ve 99 ülkeyi kapsayan veriler kullanılmıştır. Analizde bağımlı değişken olarak 2024 yılı beş yaş altı ölüm oranı kullanılmış, bağımsız değişkenler olarak ise temel içme suyu hizmetlerinden yararlanma oranı, DPT, kızamık ve Hepatit B3 aşılama kapsamı, yenidoğanlarda tetanosa karşı koruma oranı, toplam nüfus, yıllık nüfus artış oranı, temiz yakıt ve pişirme teknolojilerine erişim, karbondioksit emisyonları, tarım arazisi oranı, orman alanı, bitkisel üretim endeksi ve hayvancılık üretim endeksi olmak üzere toplam 13 gösterge değerlendirilmiştir. Araştırma sürecinde öncelikle veri setinin istatistiksel özellikleri incelenmiş, normallik testleri, korelasyon analizleri ve çoklu doğrusal bağlantı değerlendirmeleri gerçekleştirilmiştir. Veri kalitesini incelemek amacıyla İzolasyon Ormanı ve Yerel Aykırı Değer Faktörü yöntemleri kullanılarak aykırı gözlem analizi yapılmıştır. Analizler sonucunda Gambia ve Kenya aykırı gözlemler olarak belirlenmiş ancak veri hatası içermemeleri ve bilgi değeri taşımaları nedeniyle veri setinde tutulmuştur. Modelleme aşamasında dört farklı veri işleme senaryosu oluşturulmuştur. Ham hedef değişken kullanılan HAM modeli, hedef değişkene logaritmik dönüşüm uygulanmış LOG modeli, özyinelemeli özellik eleme yönteminin kullanıldığı RFE modeli ve logaritmik dönüşüm ile özellik seçiminin birlikte uygulandığı LOG+RFE modeli. Her bir senaryoda doğrusal modeller, Kernel ve komşuluk tabanlı modeller, ağaç tabanlı yöntemler, güçlendirme algoritmaları, sinir ağı modelleri ve topluluk öğrenmesi yaklaşımları olmak üzere toplam on dokuz farklı makine öğrenmesi algoritması uygulanmıştır. Modellerin hiperparametre optimizasyonu ve performans değerlendirmesi 5×10 tekrarlamalı K-katlı çapraz doğrulama yöntemi ile gerçekleştirilmiş, en iyi model sonuçlarının istatistiksel sağlamlığı bootstrap güven aralıkları kullanılarak incelenmiştir. Elde edilen bulgular, veri dönüşümü ve özellik seçimi uygulamalarının model performansını önemli ölçüde etkileyebildiğini göstermiştir. Özellikle logaritmik dönüşüm uygulanan modellerin çapraz doğrulama performanslarında iyileşme sağladığı gözlemlenmiştir. Çalışmanın metodolojik açıdan en önemli katkılarından biri, ham hedef değişken ve logaritmik dönüşüm uygulanmış hedef değişken için bağımsız özellik seçimi süreçlerinin yürütülmesi ve özellik seçiminin hedef dönüşümüne duyarlılığının incelenmesidir. Ayrıca model sonuçlarını yorumlamak amacıyla PFI, PDP, SHAP ve LIME yöntemleri kullanılmıştır. Bu analizler sayesinde yalnızca yüksek performanslı tahmin modelleri geliştirilmemiş, aynı zamanda değişkenlerin beş yaş altı ölüm oranı üzerindeki etkilerinin yönü ve büyüklüğü de ayrıntılı olarak ortaya konulmuştur. Sonuç olarak çalışma, hastalık önleme politikalarının değerlendirilmesinde makine öğrenmesi ve açıklanabilir yapay zekâ yöntemlerinin birlikte kullanılmasının sağlık politikaları açısından önemli katkılar sağlayabileceğini göstermektedir.
Özet (Çeviri)
Under five mortality rate is widely recognized as one of the most important indicators of public health, reflecting the effectiveness of healthcare systems, socioeconomic development, and environmental conditions within a country. Despite substantial improvements in global child survival over recent decades, millions of children still die before reaching their fifth birthday, particularly in low- and middle-income countries. According to international organizations such as the World Health Organization (WHO), the United Nations Children's Fund (UNICEF), and the World Bank, reducing under five mortality remains a major global health priority. The third Sustainable Development Goal aims to reduce under five mortality to fewer than 25 deaths per 1,000 live births in every country by 2030. Understanding the determinants of child mortality is therefore essential for designing effective disease prevention and public health policies. Previous studies have identified numerous factors associated with child mortality, including access to healthcare services, vaccination coverage, sanitation, safe drinking water, maternal education, nutrition, environmental quality, and socioeconomic development. However, most existing studies focus on specific regions, use limited sets of predictors, evaluate only a small number of machine learning models, and rarely investigate the effects of data preprocessing procedures such as feature selection and target transformation. Furthermore, model interpretability is often neglected despite its importance for evidence-based policymaking. The present study addresses these limitations by examining the relationships between under five mortality and a comprehensive set of health, environmental, demographic, and socioeconomic indicators using multiple machine learning approaches and explainable artificial intelligence techniques. The dataset used in this study was obtained from the World Bank database and includes information from 99 countries. The dependent variable is the under five mortality rate for 2024. Thirteen indicators associated with disease prevention and public health capacity were selected as independent variables. These indicators include access to basic drinking water services, DPT immunization coverage, measles immunization coverage, Hepatitis B immunization coverage, protection at birth against neonatal tetanus, total population, annual population growth rate, access to clean fuels and cooking technologies, carbon dioxide emissions, agricultural land ratio, forest area, crop production index, and livestock production index. The study began with a comprehensive exploratory data analysis. Normality tests, correlation analyses, and multicollinearity assessments were conducted to evaluate the statistical characteristics of the variables. To investigate data quality and identify potential anomalies, Isolation Forest (IF) and Local Outlier Factor (LOF) algorithms were employed. The analyses identified Gambia and Kenya as outlier observations. Since these observations represented valid country-level characteristics rather than measurement errors, they were retained in the dataset. Four modeling scenarios were constructed to evaluate the effects of data preprocessing on predictive performance. These included a RAW model using the original target variable, a LOG model using a logarithmically transformed target variable, an RFE model applying Recursive Feature Elimination to the original target variable, and a LOG+RFE model combining logarithmic transformation with Recursive Feature Elimination. This framework enabled a systematic comparison of the effects of target transformation and feature selection on machine learning performance. A diverse set of machine learning algorithms representing different methodological families was implemented. These included linear models (Ridge, Lasso, Elastic Net, Bayesian Ridge, Quantile Regression), kernel and neighborhood-based methods (Support Vector Regression, K-Nearest Neighbors, Gaussian Processes), tree-based algorithms (Decision Tree, Random Forest, Extra Trees, Quantile Forest), boosting methods (XGBoost, LightGBM, CatBoost, AdaBoost, Quantile Gradient Boosting), neural network models (Multilayer Perceptron and Deep Neural Networks), and a stacking ensemble model. Hyperparameter optimization and model evaluation were performed using repeated 5×10-fold cross validation. To assess statistical robustness and uncertainty, bootstrap confidence intervals were additionally calculated for model performance metrics. Although machine learning models often provide strong predictive performance, their internal decision-making mechanisms are frequently difficult to interpret. This challenge is particularly relevant for public health applications, where policymakers require transparent evidence regarding the factors influencing model predictions. To enhance model interpretability, four complementary explainable artificial intelligence techniques were employed, namely Permutation Feature Importance (PFI) for quantifying variable importance, Partial Dependence Plots (PDP) for visualizing marginal effects, SHapley Additive exPlanations (SHAP) for analyzing feature contributions and interactions, and Local Interpretable Model Agnostic Explanations (LIME) for explaining individual predictions. The combined use of these techniques enabled a comprehensive understanding of both global and local model behavior. The findings demonstrate that preprocessing strategies substantially influence machine learning performance. In particular, logarithmic transformation of the target variable generally improved cross-validation performance and enhanced model stability. The results indicate that predictive performance can vary considerably depending on the interaction between data transformation procedures and machine learning algorithms. One of the most important methodological contributions of this study is the implementation of independent feature selection procedures for the original target variable and the log-transformed target variable. This approach directly addresses the issue of feature-selection sensitivity to target transformation, a topic that has received limited attention in previous machine learning research. The explainability analyses further revealed that multiple health-related and environmental indicators contribute significantly to variations in under five mortality rates across countries. Variables associated with vaccination coverage, access to safe drinking water, and clean energy technologies consistently emerged as influential predictors. Socioeconomic and environmental indicators also exhibited meaningful contributions, highlighting the multidimensional nature of child mortality determinants. The integration of PFI, PDP, SHAP, and LIME provided complementary insights into the mechanisms underlying model predictions. While PFI identified globally important variables, PDP illustrated nonlinear relationships, SHAP quantified feature contributions and interactions, and LIME explained individual country-level predictions. Together, these methods enhanced transparency and strengthened the practical relevance of the machine learning models. This study presents a comprehensive machine learning framework for analyzing under five mortality using a multidimensional set of disease-prevention indicators from 99 countries. By systematically comparing data transformation strategies, feature selection procedures, and diverse machine learning algorithms, the research provides important methodological and practical insights. The findings demonstrate that target transformation and feature selection can significantly influence predictive performance and should be carefully considered during model development. Furthermore, the study highlights the value of integrating explainable artificial intelligence techniques into public health research to improve transparency, interpretability, and policy relevance. The proposed framework contributes to the literature by combining extensive preprocessing comparisons, multiple machine learning families, bootstrap-based uncertainty assessment, and comprehensive explainability analyses within a single analytical structure. Consequently, the study offers both methodological advancements and evidence-based guidance for policymakers seeking to reduce under five mortality and strengthen disease-prevention strategies worldwide.
Benzer Tezler
- Serological investigation of peste des petits ruminants in lambs in Iraq-Kirkuk region
Irak–Kerkük bölgesinde kuzularda küçük ruminant vebası (pestedes petits ruminants ppr)'ın seroprevalansı
SARWAT KHORSHED RAHEEM
Yüksek Lisans
İngilizce
2022
Sağlık YönetimiVan Yüzüncü Yıl ÜniversitesiSağlık Bilimleri Ana Bilim Dalı
PROF. DR. SÜLEYMAN KOZAT
- Targeting circulating breast cancer stem cells via mTOR and WNT inhibitors after bone metastasis
Sirküle meme kanseri kök hücrelerinin kemik metastazı sonrası mTOR ve WNT inhibitörleri ile hedeflenmesi̇
ÖZLEM ALTUNDAĞ ERDOĞAN
Doktora
İngilizce
2025
BiyokimyaHacettepe ÜniversitesiKök Hücre Ana Bilim Dalı
PROF. DR. BETÜL ÇELEBİ SALTIK
- Meme kanseri tanısı alan olgular ve barsak mikrobiyotası ile ilişkisinin prospektif değerlendirilmesi
Başlık çevirisi yok
MEHMET FATİH ÖZSARAY
Tıpta Uzmanlık
Türkçe
2024
Genel CerrahiKocaeli ÜniversitesiGenel Cerrahi Ana Bilim Dalı
PROF. DR. NUH ZAFER CANTÜRK
- Bipolar bozukluk tanılı hastalarda erken relaps ile ilişkili klinik ve sosyodemografik özelliklerin belirlenmesi
Determination of clinical and sociodemographic characteristics associated with early relapse in patients diagnosed with bipolar disorder
ZELİHA NUR METİN
Tıpta Uzmanlık
Türkçe
2023
PsikiyatriSağlık Bilimleri ÜniversitesiRuh Sağlığı ve Psikiyatri Ana Bilim Dalı
PROF. DR. ÖZCAN UZUN
- Çocuk yoğun bakım ünitesinde 2009-2019 yılları arasında takip edilen hastalardaki sağlık hizmeti ilişkili enfeksiyonların değerlendirilmesi
Evaluation of healthcare asssociated infections in patients followed in the pediatric intensive care unit between 2009-2019
ECE EKER
Tıpta Uzmanlık
Türkçe
2021
Çocuk Sağlığı ve HastalıklarıSağlık Bilimleri ÜniversitesiÇocuk Sağlığı ve Hastalıkları Ana Bilim Dalı
DOÇ. DR. NACİYE GÖNÜL TANIR