Geri Dön

Veri anonimleştirme tekniklerinin makine öğrenmesi ile sınıflandırma performansı üzerindeki etkilerinin karşılaştırmalı analizi

Comparative analysis of the effects of data anonymization techniques on classification performance i̇n machine learning

  1. Tez No: 1018234
  2. Yazar: ZEHRA BEGÜM AKTOLGA
  3. Danışmanlar: PROF. DR. İSMAİL HAKKI CEDİMOĞLU
  4. Tez Türü: Yüksek Lisans
  5. Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
  6. Anahtar Kelimeler: Belirtilmemiş.
  7. Yıl: 2026
  8. Dil: Türkçe
  9. Üniversite: Sakarya Üniversitesi
  10. Enstitü: Fen Bilimleri Enstitüsü
  11. Ana Bilim Dalı: Bilişim Sistemleri Mühendisliği Ana Bilim Dalı
  12. Bilim Dalı: Belirtilmemiş.
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Bu tez çalışmasında, veri anonimleştirme tekniklerinin makine öğrenmesi modellerinin performansı üzerindeki etkileri, gizlilik-kullanılabilirlik (privacy-utility trade-off) dengesi çerçevesinde deneysel olarak incelenmiştir. Çalışmanın temel amacı, kişisel verilerin korunmasına yönelik anonimleştirme uygulamalarının, farklı algoritmik yapılara sahip makine öğrenmesi modellerinde nasıl ve ne ölçüde performans değişimlerine yol açtığını nicel ve istatistiksel olarak ortaya koymaktır. Bu bağlamda hem teknik hem de hukuki ve etik boyutlar bütüncül bir yaklaşımla ele alınmıştır. Araştırmada, literatürde yaygın olarak kullanılan Adult Census Income veri seti üzerinde genelleştirme, bastırma ve türetme temelli anonimleştirme teknikleri uygulanmıştır. Orijinal ve anonimleştirilmiş veri setleri kullanılarak Lojistik Regresyon, Random Forest ve Multilayer Perceptron (MLP) modelleri eğitilmiş; modellerin performansları Accuracy, Precision, Recall, F1-skoru, ROC-AUC ve PR-AUC metrikleri üzerinden karşılaştırılmıştır. Deneysel tasarım, veri sızıntısını tamamen engelleyecek şekilde kurgulanmış; tüm ön işleme ve eşik optimizasyonu adımları yalnızca eğitim verisi üzerinde gerçekleştirilmiştir. Elde edilen sonuçların istatistiksel anlamlılığı McNemar testi ile değerlendirilmiştir. Deneysel bulgular, anonimleştirme sonrasında tüm modellerde performans metriklerinde istatistiksel olarak anlamlı ancak sınırlı düzeyde düşüşler meydana geldiğini göstermektedir. Random Forest modeli, veri ayrıntısına olan yüksek bağımlılığı nedeniyle anonimleştirmeden en fazla etkilenen yöntem olurken; MLP modeli orta düzeyde bir performans kaybı sergilemiştir. Lojistik Regresyon modeli ise parametrik yapısı sayesinde anonimleştirilmiş veri üzerinde en dayanıklı performansı göstermiştir. Buna karşın tüm modellerde ROC-AUC değerlerinin yüksek seviyelerde korunması, anonimleştirilmiş verinin belirli analitik ve karar destek senaryolarında kullanılabilirliğini sürdürebildiğine işaret etmektedir. Çalışma sonuçları, gizlilik gereksinimleri arttıkça belirli bir performans kaybının kaçınılmaz olduğunu ortaya koymakta; ancak uygun anonimleştirme düzeyi ve model seçimi ile bu kaybın kabul edilebilir sınırlar içinde tutulabileceğini göstermektedir. Bu yönüyle tez, veri gizliliği, makine öğrenmesi ve yapay zekâ uygulamalarında hem akademik literatüre deneysel kanıtlar sunmakta hem de kamu, sağlık ve finans gibi gizlilik hassasiyeti yüksek alanlarda uygulanabilir pratik çıkarımlar sağlamaktadır.

Özet (Çeviri)

In recent years, the rapid expansion of digital technologies and artificial intelligence applications has significantly increased the volume and importance of data in decision-making processes. Machine learning systems rely heavily on large-scale datasets to identify patterns, generate predictions, and support data-driven decision-making. However, many of these datasets contain sensitive personal information, which raises serious concerns regarding privacy protection and responsible data usage. As organizations increasingly rely on data analytics and artificial intelligence for strategic and operational decisions, protecting personal data has become a fundamental requirement from both legal and ethical perspectives. Regulations such as the General Data Protection Regulation (GDPR) in the European Union and the Turkish Personal Data Protection Law (KVKK) emphasize the necessity of implementing appropriate technical and organizational measures to ensure the protection of personal data. Within this context, the challenge of balancing data privacy with analytical usefulness has become a critical issue for researchers and practitioners working with data-intensive systems. This challenge is commonly described in the literature as the privacy–utility trade-off, which refers to the tension between protecting individual privacy and maintaining the analytical value of datasets. On the one hand, stronger privacy protection mechanisms reduce the risk of re-identification and misuse of personal data. On the other hand, excessive privacy protection may distort the statistical structure of datasets and reduce their usefulness for machine learning tasks. Therefore, identifying an appropriate balance between privacy protection and analytical utility is essential for the sustainable development of data-driven technologies. Data anonymization techniques are among the most widely used approaches to address this challenge. These techniques aim to transform datasets in a way that prevents the identification of individuals while preserving the information necessary for analytical tasks. However, anonymization processes often introduce a certain level of information loss, which may influence the predictive performance of machine learning models. Understanding how anonymization affects machine learning performance is therefore crucial for designing reliable and privacy-compliant artificial intelligence systems. The main objective of this thesis is to experimentally investigate the effects of data anonymization techniques on the performance of machine learning models. More specifically, the study aims to quantitatively and statistically evaluate how anonymization methods applied to protect personal data influence the predictive capabilities of machine learning algorithms with different structural characteristics. In addition to the technical analysis, the study also considers the broader legal and ethical dimensions of privacy-preserving data processing, thereby adopting a holistic perspective on the relationship between data privacy and artificial intelligence. The experimental analysis in this study is conducted using the Adult Census Income dataset, which is widely used in machine learning research for classification tasks. This dataset contains demographic and socioeconomic attributes that are commonly used to predict income levels. Within the scope of the research, several classical anonymization techniques are applied to the dataset, including generalization, suppression, and derivation-based transformations. Generalization involves replacing specific attribute values with broader categories in order to reduce identifiability, while suppression refers to masking or removing certain attribute values that may increase the risk of re-identification. Derivation-based transformations involve generating new variables derived from existing attributes in order to reduce the exposure of sensitive information while preserving analytical value. To evaluate the impact of anonymization on machine learning performance, three algorithms representing different modeling paradigms are selected: Logistic Regression, Random Forest, and Multilayer Perceptron (MLP). Logistic Regression represents a classical linear model that is widely used due to its interpretability and simplicity. Random Forest represents an ensemble-based decision tree model capable of capturing complex nonlinear relationships between variables. Multilayer Perceptron represents a neural network-based approach capable of modeling complex patterns and interactions within the data. By selecting models with different algorithmic structures, the study aims to provide a comprehensive comparison of how anonymization affects various types of machine learning methods. The models are trained using both the original dataset and the anonymized versions of the dataset. Model performances are evaluated using several widely accepted evaluation metrics, including Accuracy, Precision, Recall, F1-score, ROC-AUC, and PR-AUC. These metrics enable a multidimensional evaluation of classification performance, allowing for a more comprehensive assessment of the effects of anonymization on predictive accuracy and classification quality. In order to ensure the reliability of the experimental results, the experimental design is constructed to prevent data leakage. All preprocessing steps, including data cleaning, encoding, scaling, and threshold optimization, are performed exclusively on the training dataset. The trained models are then evaluated on separate test datasets to obtain unbiased performance estimates. In addition, the statistical significance of the observed performance differences is evaluated using the McNemar statistical test. The experimental results indicate that anonymization techniques lead to statistically significant but limited decreases in model performance across all evaluated algorithms. Among the examined models, the Random Forest algorithm is the most affected by anonymization. This result can be explained by the strong dependence of tree-based models on detailed feature representations and precise splitting criteria. When anonymization techniques generalize or suppress certain attributes, the ability of decision trees to identify optimal split points may be reduced. The Multilayer Perceptron (MLP) model exhibits a moderate level of performance degradation. Although neural networks are capable of learning complex relationships within the data, anonymization may still disrupt important correlations between variables, which can influence the learning process. In contrast, Logistic Regression demonstrates the highest robustness on anonymized data. Due to its parametric structure and relatively simple decision boundaries, the model appears to be less sensitive to certain types of information loss introduced by anonymization. Despite the observed decreases in some performance metrics, the results also show that relatively high ROC-AUC values are preserved across all models. This finding suggests that anonymized datasets can still retain meaningful predictive capability under certain conditions. In particular, anonymized data may remain useful for decision-support systems and predictive analytics in scenarios where strict privacy protection is required. Overall, the findings confirm that increasing privacy protection through anonymization inevitably introduces a certain degree of information loss. However, the results also demonstrate that these performance losses can remain within acceptable limits when appropriate anonymization strategies and model selections are applied. Therefore, it is possible to achieve a practical balance between data privacy and analytical utility. In this respect, the thesis contributes empirical evidence to the literature on privacy-preserving machine learning and provides practical insights for the design of artificial intelligence systems in privacy-sensitive domains such as healthcare, finance, and public administration. Future studies may further extend this research by evaluating additional anonymization techniques, applying the methodology to different datasets, and examining the interaction between anonymization and more advanced machine learning architectures.

Benzer Tezler

  1. Veri anonimleştirme ve sentetik veri üretimi tekniklerinin karşılaştırılması

    Comparative analysis of data anonymization and synthetic data generation techniques

    KAAN KIVIRCIK

    Yüksek Lisans

    Türkçe

    Türkçe

    2025

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolEge Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    DR. ÖĞR. ÜYESİ EMİNE SEZER

    PROF. DR. DENİZ KILINÇ

  2. Preserving privacy of health data residing in HL7 FHIR repositories through de-identification

    HL7 FHIR kaynaklarında bulunan sağlık verilerinin gizliliğinin kimliksizleştirme yoluyla korunması

    EZELSU ŞİMŞEK YILGIN

    Yüksek Lisans

    İngilizce

    İngilizce

    2022

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolOrta Doğu Teknik Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    PROF. DR. PINAR KARAGÖZ

  3. Towards differentially private data publishing for effective privacy research

    Akıllı sayaç verilerinde etkıi mahremiyet araştırmaları için diferansiyel gizli gürültü ekleme

    MOHAMED MEDHAT MOHAMED ALI ZEINA

    Yüksek Lisans

    İngilizce

    İngilizce

    2023

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolSabancı Üniversitesi

    Bilgisayar Bilimleri ve Mühendisliği Ana Bilim Dalı

    PROF. DR. ALBERT LEVİ

  4. The bid/no-bid framework utilizing Node2Vec for international contractors

    Uluslararası yükleniciler için Node2Vec ile teklif/red çerçevesi oluşturulması

    HALİT FERDİ AVCI

    Yüksek Lisans

    İngilizce

    İngilizce

    2024

    İnşaat Mühendisliğiİstanbul Teknik Üniversitesi

    İnşaat Mühendisliği Ana Bilim Dalı

    DOÇ. DR. ONUR BEHZAT TOKDEMİR

  5. Veri sızıntısı önlemede veri maskeleme yöntemlerinin etkinliği

    The effectiveness of data masking techniques in data loss prevention

    YEŞİM ALAN KILIÇ

    Yüksek Lisans

    Türkçe

    Türkçe

    2024

    Bilgi ve Belge YönetimiBahçeşehir Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    DR. ÖĞR. ÜYESİ AHMET NACİ ÜNAL