Geri Dön

Saldırı tespit sistemlerinde yapay zekâ tabanlı anomali tespiti

Artificial intelligence-based anomaly detection in intrusion detection systems

  1. Tez No: 1015641
  2. Yazar: AWAB DAW SAAD KHALEFA
  3. Danışmanlar: PROF. DR. AHMET ZENGİN
  4. Tez Türü: Yüksek Lisans
  5. Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
  6. Anahtar Kelimeler: Belirtilmemiş.
  7. Yıl: 2025
  8. Dil: Türkçe
  9. Üniversite: Sakarya Üniversitesi
  10. Enstitü: Fen Bilimleri Enstitüsü
  11. Ana Bilim Dalı: Bilgisayar Mühendisliği Ana Bilim Dalı
  12. Bilim Dalı: Belirtilmemiş.
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Bu tez çalışması, Yapay Zekâ (YZ) ve Derin Öğrenme (DÖ) tekniklerinin Saldırı Tespit Sistemleri (STS) içerisindeki etkinliğini artırmaya yönelik bir yaklaşım sunmaktadır. Günümüzde dijital altyapıların hızla genişlemesi, ağ tabanlı sistemlerin karmaşıklığını artırmakta ve bu durum siber tehditlerin daha çeşitli ve gelişmiş hâle gelmesine yol açmaktadır. Geleneksel imza tabanlı saldırı tespit yöntemleri bilinen tehditleri tespit etmede etkili olsa da, yeni ortaya çıkan ve daha önce tanımlanmamış saldırı türlerini belirlemede yetersiz kalabilmektedir. Bu nedenle, bilinmeyen ve sıfırıncı gün saldırılarını da tanıyabilen daha esnek ve öğrenebilir sistemlere ihtiyaç duyulmaktadır. Bu çalışmanın temel amacı, geleneksel imza tabanlı tespit yöntemlerinin ötesine geçerek, bilinmeyen ve sıfırıncı gün saldırılarını da tanıyabilen anomali tabanlı bir sistem geliştirmektir. Bu amaç doğrultusunda, Gürültü Azaltan Otomatik Kodlayıcı (GAOK) mimarisi temel alınmış ve denetimsiz öğrenme yaklaşımıyla ağ trafiğindeki normal davranış modelleri öğrenilmiştir. GAOK tabanlı model yalnızca normal ağ trafiği verileriyle eğitilmiş, böylece sistemin ağın normal işleyişini otomatik olarak öğrenmesi sağlanmıştır. Bu yaklaşım sayesinde model, ağ trafiğinin temel yapısal özelliklerini öğrenerek normal davranış kalıplarını temsil eden bir profil oluşturmuştur. Test aşamasında model, gelen ağ verileri için yeniden yapılandırma hatasını hesaplamakta ve belirlenen eşik değerini aşan örnekleri potansiyel saldırı olarak sınıflandırmaktadır. Bu yöntem, saldırıların önceden tanımlanmış bir imzaya bağlı kalmadan tespit edilmesini mümkün kılarak, sistemin dinamik ve öngörülemeyen tehditlere karşı da etkili olmasını sağlamaktadır. Böylece önerilen yaklaşım, ağ güvenliği sistemlerinde adaptif ve veri odaklı bir tespit mekanizması sunmaktadır. Çalışma kapsamında modelin performansı, üç farklı ve yaygın olarak kullanılan ağ güvenliği veri seti üzerinde test edilmiştir: NSL-KDD, UNSW-NB15 ve CTU-IoT (IoT-23). Bu veri setleri, klasik ağ trafiğinden modern ağ saldırılarına ve IoT tabanlı sistemlere kadar geniş bir saldırı yelpazesini temsil etmektedir. Veri ön işleme aşamasında, kategorik özellikler için tek-sıcak kodlama yöntemi uygulanmış, sayısal değişkenler ise Min-Max normalizasyonu kullanılarak ölçeklendirilmiştir. Ayrıca modelin daha sağlam temsiller öğrenebilmesi ve gürültüye karşı dayanıklılığının artırılması amacıyla eğitim sürecinde seyreltme katmanları kullanılmıştır. Bu sayede modelin aşırı öğrenme riski azaltılmış ve genelleme yeteneği güçlendirilmiştir. Elde edilen sonuçlar, geliştirilen modelin özellikle NSL-KDD veri setinde %86 doğruluk ve %87 F1-skoru ile güçlü bir performans sergilediğini göstermektedir. UNSW-NB15 veri setinde %94 doğruluk elde edilmesine rağmen, belirli saldırı türlerinde yanlış pozitif oranlarının arttığı gözlemlenmiştir. IoT tabanlı CTU-IoT veri setinde ise, modelin başarısının büyük ölçüde eşik değeri seçimine bağlı olduğu ve farklı senaryolarda bu değerin sistemin hassasiyetini doğrudan etkilediği belirlenmiştir. Bu bulgular, anomali tabanlı sistemlerde eşik belirleme stratejisinin model performansı üzerinde kritik bir rol oynadığını göstermektedir. Çalışma sonucunda, YZ tabanlı otomatik kodlayıcıların ağ güvenliğinde anomali tespiti için güçlü ve esnek bir alternatif sunduğu ortaya konmuştur. Bununla birlikte, verilerin dengesiz dağılımı, nadir saldırı türlerinin sınırlı temsili ve eşik belirleme zorlukları performans üzerinde etkili olan önemli faktörler olarak değerlendirilmiştir. Bu nedenle, gelecekteki çalışmalarda hibrit tespit sistemlerinin kullanımı, uyarlamalı eşikleme tekniklerinin geliştirilmesi ve öznitelik seçimi yöntemlerinin optimize edilmesi önerilmektedir. Ayrıca farklı derin öğrenme mimarilerinin ve açıklanabilir yapay zekâ tekniklerinin kullanılması, model kararlarının daha şeffaf hâle getirilmesine katkı sağlayabilir. Sonuç olarak bu tez, derin öğrenme yöntemlerinin saldırı tespit sistemlerine entegrasyonu konusunda anlamlı bir katkı sunmakta; özellikle GAOK tabanlı modellerin farklı ağ ortamlarında yüksek doğruluk oranlarıyla anomalileri tespit edebildiğini göstermektedir. Önerilen yaklaşım, hem geleneksel ağ altyapılarında hem de IoT tabanlı ortamlarda uygulanabilir bir çözüm sunarak gelecekte yapay zekâ destekli siber güvenlik sistemlerinin geliştirilmesinde önemli bir temel oluşturma potansiyeline sahiptir.

Özet (Çeviri)

The rapid growth of digital technologies and network-connected infrastructures has significantly increased the importance of cybersecurity in modern society. Critical infrastructures such as transportation systems, financial networks, healthcare services, industrial control systems, and communication networks rely heavily on interconnected digital environments. While these technologies improve efficiency and connectivity, they also expand the attack surface and create new vulnerabilities that can be exploited by malicious actors. As cyber threats continue to evolve in complexity and sophistication, protecting network systems from unauthorized access and malicious activities has become a fundamental challenge. Intrusion Detection Systems (IDS) play a critical role in maintaining the security of computer networks. These systems monitor network traffic and system activities in order to detect suspicious behavior, policy violations, and potential cyberattacks. Traditionally, IDS solutions have relied on signature-based detection techniques, which identify attacks by matching network patterns with known attack signatures stored in a database. Although this approach is highly effective in detecting previously known attacks, it has a major limitation: it cannot detect new or previously unseen attacks, commonly referred to as zero-day attacks. To overcome this limitation, anomaly-based intrusion detection techniques have been proposed. Instead of relying on predefined attack signatures, anomaly-based systems learn the normal behavior of network traffic and detect deviations from this learned pattern as potential threats. These systems have the advantage of detecting unknown attacks; however, they often suffer from high false-positive rates, especially in complex and dynamic network environments. Recent advancements in artificial intelligence and deep learning have provided new opportunities to improve anomaly-based intrusion detection systems. Deep learning models are capable of automatically learning complex patterns and relationships from large-scale datasets, making them particularly suitable for network traffic analysis. Among these techniques, autoencoders have gained considerable attention due to their ability to learn compressed representations of input data and reconstruct it with minimal error. This thesis proposes an anomaly-based intrusion detection model based on a Denoising Autoencoder (DAE). The proposed approach utilizes an unsupervised learning strategy in which the model is trained using only normal network traffic data. The purpose of this training strategy is to enable the model to learn the intrinsic characteristics of legitimate network behavior. During the training process, artificial noise is added to the input data, allowing the DAE to learn how to reconstruct clean data from corrupted inputs. This process enhances the model's robustness and improves its generalization capability when encountering noisy or partially corrupted network traffic. Once the model has been trained, anomaly detection is performed by calculating the reconstruction error between the original input and the reconstructed output. If the reconstruction error exceeds a predefined threshold, the data instance is classified as anomalous and potentially malicious. This approach allows the model to detect abnormal patterns in network traffic without requiring labeled attack data. To evaluate the effectiveness of the proposed intrusion detection model, experiments were conducted using three widely recognized cybersecurity datasets: NSL-KDD, UNSW-NB15, and CTU-IoT. These datasets were selected because they represent different types of network environments and attack scenarios. The NSL-KDD dataset represents traditional network traffic and is commonly used in IDS research. The UNSW-NB15 dataset contains modern attack scenarios and more realistic network traffic patterns. The CTU-IoT dataset focuses on Internet of Things (IoT) environments, which introduce additional security challenges due to the large number of heterogeneous and resource-constrained devices. Before training the model, a comprehensive data preprocessing pipeline was implemented. This process included data cleaning, feature selection, categorical feature encoding, normalization of numerical features, and optimization of data types to reduce computational overhead. Proper preprocessing is essential in machine learning-based intrusion detection systems because raw network data often contains redundant, noisy, or highly imbalanced features that may negatively impact model performance. The architecture of the proposed Denoising Autoencoder consists of an encoder, a bottleneck layer, and a decoder. The encoder compresses the input data into a lower-dimensional latent representation that captures the essential characteristics of normal network behavior. The bottleneck layer represents the most compact representation of the input data and forces the model to learn meaningful features. The decoder then reconstructs the input data from this compressed representation. Experimental results demonstrate that the proposed model is capable of effectively identifying anomalies in network traffic across different datasets. On the NSL-KDD dataset, the model achieved approximately 86% accuracy and an F1-score of around 87%, indicating strong detection performance. For the UNSW-NB15 dataset, the model achieved approximately 94% accuracy, although certain attack categories produced higher false-positive rates due to their similarity with normal traffic patterns. In the CTU-IoT dataset, the performance of the model was found to be sensitive to the selected anomaly threshold, highlighting the importance of threshold optimization in anomaly-based detection systems. Despite the promising results obtained in this study, several limitations were identified during the experimental analysis. One of the primary challenges is the imbalance of attack classes in the datasets. Many cybersecurity datasets contain a significantly larger number of normal samples compared to certain rare attack types. This imbalance can cause the model to become biased toward dominant classes, reducing its ability to detect minority attacks. Additionally, the overlap between normal and malicious traffic features may lead to false-positive detections in some scenarios. Another challenge involves the selection of an appropriate anomaly threshold. Since anomaly detection relies on reconstruction error values, determining the optimal threshold is crucial for balancing false-positive and false-negative rates. A fixed threshold may not perform equally well across different datasets or network environments, suggesting the need for adaptive threshold mechanisms. To address these limitations, several potential improvements are suggested for future research. Hybrid intrusion detection approaches that combine anomaly-based detection with signature-based methods could improve overall detection accuracy. Additionally, advanced feature selection techniques and dimensionality reduction methods may enhance the model's ability to distinguish between normal and malicious traffic patterns. Future work may also explore more advanced deep learning architectures, such as Variational Autoencoders (VAE), Generative Adversarial Networks (GAN), or hybrid deep learning models that integrate supervised learning components. Furthermore, the integration of Explainable Artificial Intelligence (XAI) techniques could improve the transparency and interpretability of deep learning-based intrusion detection systems, making them more suitable for deployment in real-world cybersecurity environments. In conclusion, this thesis demonstrates that Denoising Autoencoders provide a powerful and flexible framework for anomaly-based intrusion detection. The proposed model successfully learns the underlying patterns of normal network behavior and detects deviations that may indicate cyber threats. Although challenges such as dataset imbalance and threshold selection remain, the findings of this research highlight the significant potential of deep learning techniques in developing intelligent and adaptive cybersecurity solutions capable of addressing emerging network security threats.

Benzer Tezler

  1. Comparison of intrusion detection for the internet of things with machine and deep learning methods

    Makine ve derin öğrenme yöntemleri ile nesnelerin interneti için saldırı tespitinin karşılaştırılması

    SIHAM AMAROUCHE

    Yüksek Lisans

    İngilizce

    İngilizce

    2021

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolKocaeli Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    PROF. DR. KEREM KÜÇÜK

  2. Anomaly detection in ınternet of medical things using deep learning

    Anomaly detect ionin internet of medical things using deep learning

    AYŞE BETÜL BÜKEN

    Yüksek Lisans

    İngilizce

    İngilizce

    2025

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolSakarya Üniversitesi

    Yazılım Mühendisliği Ana Bilim Dalı

    PROF. DR. DEVRİM AKGÜN

  3. The role of ai and machine learning in cloud intrusiondetection systems

    Bulut saldırı tespit sistemlerinde yapay zeka ve makineöğreniminin rolü

    ANAS CHEHADEH

    Yüksek Lisans

    İngilizce

    İngilizce

    2026

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolBahçeşehir Üniversitesi

    Siber Güvenlik Ana Bilim Dalı

    DOÇ. DR. YÜCEL BATU SALMAN

  4. DDOS saldırılarının tespit edilmesinde makine öğrenimi yöntemlerinin uygulanması

    Application of machine learning methods in determining DDOSattacks

    TUĞBA AYTAÇ

    Yüksek Lisans

    Türkçe

    Türkçe

    2020

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Ticaret Üniversitesi

    Bilgisayar Mühendisliği Ana Bilim Dalı

    PROF. DR. ABDÜL HALİM ZAİM

    DOÇ. DR. MUHAMMED ALİ AYDIN

  5. Ağ trafiğinin analizi, anomali tespiti ve değerlendirme

    Analysis of network traffic, anomaly detection and evaluation

    AKIN ASLAN

    Yüksek Lisans

    Türkçe

    Türkçe

    2017

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik Üniversitesi

    Bilişim Uygulamaları Ana Bilim Dalı

    DOÇ. DR. ENVER ÖZDEMİR