Machine learning for big data in financial fraud detection
Finansal dolandırıcılık tespitinde büyük veriler için makine öğrenimi
- Tez No: 1023479
- Danışmanlar: PROF. DR. MUSTAFA ÖZBAYRAK
- Tez Türü: Yüksek Lisans
- Konular: Yönetim Bilişim Sistemleri, Management Information Systems
- Anahtar Kelimeler: Belirtilmemiş.
- Yıl: 2026
- Dil: İngilizce
- Üniversite: Bahçeşehir Üniversitesi
- Enstitü: Lisansüstü Eğitim Enstitüsü
- Ana Bilim Dalı: Büyük Veri Analitiği ve Yönetimi Ana Bilim Dalı
- Bilim Dalı: Büyük Veri Analitiği ve Yönetimi Bilim Dalı
- Sayfa Sayısı: Belirtilmemiş.
Özet
Bu tez, finansal işlemlerde dolandırıcılığın tespit edilmesine yönelik ölçeklenebilir, yüksek performanslı ve açıklanabilir bir makine öğrenmesi yaklaşımı geliştirmeyi amaçlamaktadır. Dijital finansal işlemlerin hızla artması, dolandırıcılık yöntemlerinin daha karmaşık ve değişken hale gelmesine neden olmuş; geleneksel kural tabanlı sistemler yüksek yanlış pozitif oranları ve düşük uyum yetenekleri nedeniyle yetersiz kalmıştır. Bu çalışma, denetimli ve denetimsiz öğrenme yöntemlerini birleştiren bütüncül bir dolandırıcılık tespit hattı önermektedir. Çalışmada, aşırı dengesiz bir veri kümesi üzerinde Lojistik Regresyon, Rastgele Orman ve XGBoost modelleri karşılaştırılmış; sınıf dengesizliğini azaltmak amacıyla SMOTE ve ADASYN yöntemleri sistematik olarak değerlendirilmiştir. Ayrıca, daha önce görülmemiş veya nadir dolandırıcılık örüntülerini tespit edebilmek için denetimsiz bir Otokodlayıcı (Autoencoder) modeli kullanılmıştır. Model performansı, dolandırıcılık tespiti açısından kritik olan kesinlik, duyarlılık, F1-skoru ve Hassasiyet–Duyarlılık Eğrisi Altındaki Alan (AUPRC) metrikleri kullanılarak değerlendirilmiştir. Karar eşikleri optimize edilmiş ve model şeffaflığını sağlamak amacıyla SHAP tabanlı açıklanabilir yapay zekâ yöntemleri uygulanmıştır. Elde edilen sonuçlar, SMOTE ile eğitilmiş XGBoost modelinin doğruluk ve duyarlılık arasında en dengeli performansı sunduğunu göstermektedir. Otokodlayıcı modelin ise sıfırıncı gün (zero-day) dolandırıcılıklarının tespitinde tamamlayıcı bir araç olarak kullanılabileceği ortaya konmuştur. Bu çalışma, gerçek dünya finansal sistemleri için uygulanabilir, yorumlanabilir ve yüksek doğruluklu bir dolandırıcılık tespit çerçevesi sunmaktadır.
Özet (Çeviri)
The expeditious expansion of digital financial transactions has not only enhanced convenience but has also broadened the sources of defraud, thus raising the issue of detecting fraud as a thorny problem to financial institutions. The systems that operate using traditional rules find it difficult to keep up with the changing fraud trends and in most cases tend to have high false positive rates and a low level of detection. The research paper builds upon a complete machine learning pipeline in detecting fraudulent financial transactions, supervised ensemble models, unsupervised anomaly detection, and explainable artificial intelligence (XAI). The pipeline assesses the Logistic Regression, Random Forest and XGBoost models on an imbalanced dataset of credit cards transactions of 284,807 transactions with a 0.172% fraud rate. The comparison of imbalance mitigation strategies SMOTE and ADASYN is done in a systematic way. Also, an Autoencoder is deployed to make unsupervised discoveries of fraudulent patterns that are rare or have never been detected before. Model evaluation aims to emphasize metrics that are of relevance to fraud, such as Precision, Recall, F1-score and Area Under the Precision-Recall Curve (AUPRC), with thresholds of decisions being optimized to balance the level of detection. SHAP explainability is used to guarantee transparency, interpretability on a feature level, and regulatory compliance. The findings show that XGBoost with SMOTE can be considered as the most favorable trade-off between accuracy and recall since it has the highest F1-score (0.8715) and good AUPRC (0.8658), and the decision behavior is stable and interpretable. Random Forest with SMOTE has the highest AUPRC but lower recall and high cost of computation. Logistic Regression has a high recall and a high number of false positives. The Autoencoder shows high anomaly detection but low precision, so it is an appropriate addition to use as a supplemental tool to detect zero-day fraud. SHAP analysis emphasizes major predictive features, which increases transparency and auditability. The work establishes the fact that contemporary machine learning methods are able to provide highly accurate, interpretable, and deployable fraud detection systems. The suggested pipeline offers a feasible structure of a real-world financial application and provides information about future studies on cost-sensitive learning, time-sensitive modeling, real-time implementation, and fraud detection that considers fairness.
Benzer Tezler
- Machine learning approach for external fraud detection
Dış saldırıların belirlenmesi için makine öğrenimi yaklaşımı
AJI MUBALAIKE
Yüksek Lisans
İngilizce
2018
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik ÜniversitesiBilişim Uygulamaları Ana Bilim Dalı
PROF. DR. ERTUĞRUL KARAÇUHA
PROF. DR. EŞREF ADALI
- İklimlendirme sistemleri üzerinde makine öğrenmesi ile anomali tespiti
Anomaly detection with machine learning on air conditioning systems
REFİK KİBAR
Yüksek Lisans
Türkçe
2023
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolSakarya ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
DR. ÖĞR. ÜYESİ MUHAMMED FATİH ADAK
DR. ÖĞR. ÜYESİ KEVSER OVAZ AKPINAR
- Digital transformation in finance exploring the evolution and impact of Fintech innovation
Finansta dijital dönüşüm Fintech inovasyonunun evrimi ve etkisini keşfedin
ALI SATTAR JABBAR AL-HRAISHAWI
Yüksek Lisans
İngilizce
2024
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolAltınbaş ÜniversitesiBilişim Teknolojileri Ana Bilim Dalı
DR. ÖĞR. ÜYESİ OGUZ KARAN
- Türk sigorta sektöründe kullanılan yeni dijital teknolojiler (Insurtech)
New digital technologies used in Turkish insurance sector (Insurtech)
PINAR AYDEMİR
Doktora
Türkçe
2023
Sigortacılıkİstanbul Ticaret ÜniversitesiSigortacılık ve Risk Yönetimi Ana Bilim Dalı
PROF. DR. ÜNAL HALİT ÖZDEN
- Hile riskinin tespitinde f-skor modeli ve hile beşgeni teorisi üzerine BIST'de yapılan bir araştırma
An investigation in BIST on f-score model and pentagon theory for the detection of fraud risk
ECE ÇEVİK