Veri anonimleştirme tekniklerinin makine öğrenmesi ile sınıflandırma performansı üzerindeki etkilerinin karşılaştırmalı analizi
Comparative analysis of the effects of data anonymization techniques on classification performance i̇n machine learning
- Tez No: 1018234
- Danışmanlar: PROF. DR. İSMAİL HAKKI CEDİMOĞLU
- Tez Türü: Yüksek Lisans
- Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
- Anahtar Kelimeler: Belirtilmemiş.
- Yıl: 2026
- Dil: Türkçe
- Üniversite: Sakarya Üniversitesi
- Enstitü: Fen Bilimleri Enstitüsü
- Ana Bilim Dalı: Bilişim Sistemleri Mühendisliği Ana Bilim Dalı
- Bilim Dalı: Belirtilmemiş.
- Sayfa Sayısı: Belirtilmemiş.
Özet
Bu tez çalışmasında, veri anonimleştirme tekniklerinin makine öğrenmesi modellerinin performansı üzerindeki etkileri, gizlilik-kullanılabilirlik (privacy-utility trade-off) dengesi çerçevesinde deneysel olarak incelenmiştir. Çalışmanın temel amacı, kişisel verilerin korunmasına yönelik anonimleştirme uygulamalarının, farklı algoritmik yapılara sahip makine öğrenmesi modellerinde nasıl ve ne ölçüde performans değişimlerine yol açtığını nicel ve istatistiksel olarak ortaya koymaktır. Bu bağlamda hem teknik hem de hukuki ve etik boyutlar bütüncül bir yaklaşımla ele alınmıştır. Araştırmada, literatürde yaygın olarak kullanılan Adult Census Income veri seti üzerinde genelleştirme, bastırma ve türetme temelli anonimleştirme teknikleri uygulanmıştır. Orijinal ve anonimleştirilmiş veri setleri kullanılarak Lojistik Regresyon, Random Forest ve Multilayer Perceptron (MLP) modelleri eğitilmiş; modellerin performansları Accuracy, Precision, Recall, F1-skoru, ROC-AUC ve PR-AUC metrikleri üzerinden karşılaştırılmıştır. Deneysel tasarım, veri sızıntısını tamamen engelleyecek şekilde kurgulanmış; tüm ön işleme ve eşik optimizasyonu adımları yalnızca eğitim verisi üzerinde gerçekleştirilmiştir. Elde edilen sonuçların istatistiksel anlamlılığı McNemar testi ile değerlendirilmiştir. Deneysel bulgular, anonimleştirme sonrasında tüm modellerde performans metriklerinde istatistiksel olarak anlamlı ancak sınırlı düzeyde düşüşler meydana geldiğini göstermektedir. Random Forest modeli, veri ayrıntısına olan yüksek bağımlılığı nedeniyle anonimleştirmeden en fazla etkilenen yöntem olurken; MLP modeli orta düzeyde bir performans kaybı sergilemiştir. Lojistik Regresyon modeli ise parametrik yapısı sayesinde anonimleştirilmiş veri üzerinde en dayanıklı performansı göstermiştir. Buna karşın tüm modellerde ROC-AUC değerlerinin yüksek seviyelerde korunması, anonimleştirilmiş verinin belirli analitik ve karar destek senaryolarında kullanılabilirliğini sürdürebildiğine işaret etmektedir. Çalışma sonuçları, gizlilik gereksinimleri arttıkça belirli bir performans kaybının kaçınılmaz olduğunu ortaya koymakta; ancak uygun anonimleştirme düzeyi ve model seçimi ile bu kaybın kabul edilebilir sınırlar içinde tutulabileceğini göstermektedir. Bu yönüyle tez, veri gizliliği, makine öğrenmesi ve yapay zekâ uygulamalarında hem akademik literatüre deneysel kanıtlar sunmakta hem de kamu, sağlık ve finans gibi gizlilik hassasiyeti yüksek alanlarda uygulanabilir pratik çıkarımlar sağlamaktadır.
Özet (Çeviri)
In recent years, the rapid expansion of digital technologies and artificial intelligence applications has significantly increased the volume and importance of data in decision-making processes. Machine learning systems rely heavily on large-scale datasets to identify patterns, generate predictions, and support data-driven decision-making. However, many of these datasets contain sensitive personal information, which raises serious concerns regarding privacy protection and responsible data usage. As organizations increasingly rely on data analytics and artificial intelligence for strategic and operational decisions, protecting personal data has become a fundamental requirement from both legal and ethical perspectives. Regulations such as the General Data Protection Regulation (GDPR) in the European Union and the Turkish Personal Data Protection Law (KVKK) emphasize the necessity of implementing appropriate technical and organizational measures to ensure the protection of personal data. Within this context, the challenge of balancing data privacy with analytical usefulness has become a critical issue for researchers and practitioners working with data-intensive systems. This challenge is commonly described in the literature as the privacy–utility trade-off, which refers to the tension between protecting individual privacy and maintaining the analytical value of datasets. On the one hand, stronger privacy protection mechanisms reduce the risk of re-identification and misuse of personal data. On the other hand, excessive privacy protection may distort the statistical structure of datasets and reduce their usefulness for machine learning tasks. Therefore, identifying an appropriate balance between privacy protection and analytical utility is essential for the sustainable development of data-driven technologies. Data anonymization techniques are among the most widely used approaches to address this challenge. These techniques aim to transform datasets in a way that prevents the identification of individuals while preserving the information necessary for analytical tasks. However, anonymization processes often introduce a certain level of information loss, which may influence the predictive performance of machine learning models. Understanding how anonymization affects machine learning performance is therefore crucial for designing reliable and privacy-compliant artificial intelligence systems. The main objective of this thesis is to experimentally investigate the effects of data anonymization techniques on the performance of machine learning models. More specifically, the study aims to quantitatively and statistically evaluate how anonymization methods applied to protect personal data influence the predictive capabilities of machine learning algorithms with different structural characteristics. In addition to the technical analysis, the study also considers the broader legal and ethical dimensions of privacy-preserving data processing, thereby adopting a holistic perspective on the relationship between data privacy and artificial intelligence. The experimental analysis in this study is conducted using the Adult Census Income dataset, which is widely used in machine learning research for classification tasks. This dataset contains demographic and socioeconomic attributes that are commonly used to predict income levels. Within the scope of the research, several classical anonymization techniques are applied to the dataset, including generalization, suppression, and derivation-based transformations. Generalization involves replacing specific attribute values with broader categories in order to reduce identifiability, while suppression refers to masking or removing certain attribute values that may increase the risk of re-identification. Derivation-based transformations involve generating new variables derived from existing attributes in order to reduce the exposure of sensitive information while preserving analytical value. To evaluate the impact of anonymization on machine learning performance, three algorithms representing different modeling paradigms are selected: Logistic Regression, Random Forest, and Multilayer Perceptron (MLP). Logistic Regression represents a classical linear model that is widely used due to its interpretability and simplicity. Random Forest represents an ensemble-based decision tree model capable of capturing complex nonlinear relationships between variables. Multilayer Perceptron represents a neural network-based approach capable of modeling complex patterns and interactions within the data. By selecting models with different algorithmic structures, the study aims to provide a comprehensive comparison of how anonymization affects various types of machine learning methods. The models are trained using both the original dataset and the anonymized versions of the dataset. Model performances are evaluated using several widely accepted evaluation metrics, including Accuracy, Precision, Recall, F1-score, ROC-AUC, and PR-AUC. These metrics enable a multidimensional evaluation of classification performance, allowing for a more comprehensive assessment of the effects of anonymization on predictive accuracy and classification quality. In order to ensure the reliability of the experimental results, the experimental design is constructed to prevent data leakage. All preprocessing steps, including data cleaning, encoding, scaling, and threshold optimization, are performed exclusively on the training dataset. The trained models are then evaluated on separate test datasets to obtain unbiased performance estimates. In addition, the statistical significance of the observed performance differences is evaluated using the McNemar statistical test. The experimental results indicate that anonymization techniques lead to statistically significant but limited decreases in model performance across all evaluated algorithms. Among the examined models, the Random Forest algorithm is the most affected by anonymization. This result can be explained by the strong dependence of tree-based models on detailed feature representations and precise splitting criteria. When anonymization techniques generalize or suppress certain attributes, the ability of decision trees to identify optimal split points may be reduced. The Multilayer Perceptron (MLP) model exhibits a moderate level of performance degradation. Although neural networks are capable of learning complex relationships within the data, anonymization may still disrupt important correlations between variables, which can influence the learning process. In contrast, Logistic Regression demonstrates the highest robustness on anonymized data. Due to its parametric structure and relatively simple decision boundaries, the model appears to be less sensitive to certain types of information loss introduced by anonymization. Despite the observed decreases in some performance metrics, the results also show that relatively high ROC-AUC values are preserved across all models. This finding suggests that anonymized datasets can still retain meaningful predictive capability under certain conditions. In particular, anonymized data may remain useful for decision-support systems and predictive analytics in scenarios where strict privacy protection is required. Overall, the findings confirm that increasing privacy protection through anonymization inevitably introduces a certain degree of information loss. However, the results also demonstrate that these performance losses can remain within acceptable limits when appropriate anonymization strategies and model selections are applied. Therefore, it is possible to achieve a practical balance between data privacy and analytical utility. In this respect, the thesis contributes empirical evidence to the literature on privacy-preserving machine learning and provides practical insights for the design of artificial intelligence systems in privacy-sensitive domains such as healthcare, finance, and public administration. Future studies may further extend this research by evaluating additional anonymization techniques, applying the methodology to different datasets, and examining the interaction between anonymization and more advanced machine learning architectures.
Benzer Tezler
- Veri anonimleştirme ve sentetik veri üretimi tekniklerinin karşılaştırılması
Comparative analysis of data anonymization and synthetic data generation techniques
KAAN KIVIRCIK
Yüksek Lisans
Türkçe
2025
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolEge ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
DR. ÖĞR. ÜYESİ EMİNE SEZER
PROF. DR. DENİZ KILINÇ
- Preserving privacy of health data residing in HL7 FHIR repositories through de-identification
HL7 FHIR kaynaklarında bulunan sağlık verilerinin gizliliğinin kimliksizleştirme yoluyla korunması
EZELSU ŞİMŞEK YILGIN
Yüksek Lisans
İngilizce
2022
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolOrta Doğu Teknik ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
PROF. DR. PINAR KARAGÖZ
- Towards differentially private data publishing for effective privacy research
Akıllı sayaç verilerinde etkıi mahremiyet araştırmaları için diferansiyel gizli gürültü ekleme
MOHAMED MEDHAT MOHAMED ALI ZEINA
Yüksek Lisans
İngilizce
2023
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolSabancı ÜniversitesiBilgisayar Bilimleri ve Mühendisliği Ana Bilim Dalı
PROF. DR. ALBERT LEVİ
- The bid/no-bid framework utilizing Node2Vec for international contractors
Uluslararası yükleniciler için Node2Vec ile teklif/red çerçevesi oluşturulması
HALİT FERDİ AVCI
Yüksek Lisans
İngilizce
2024
İnşaat Mühendisliğiİstanbul Teknik Üniversitesiİnşaat Mühendisliği Ana Bilim Dalı
DOÇ. DR. ONUR BEHZAT TOKDEMİR
- Veri sızıntısı önlemede veri maskeleme yöntemlerinin etkinliği
The effectiveness of data masking techniques in data loss prevention
YEŞİM ALAN KILIÇ
Yüksek Lisans
Türkçe
2024
Bilgi ve Belge YönetimiBahçeşehir ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
DR. ÖĞR. ÜYESİ AHMET NACİ ÜNAL