Derin öğrenme modellerinin transfer öğrenme ile farklı veri setleri üzerindeki performanslarının karşılaştırılması
Comparative performance analysis of deep learning models on different datasets using transfer learning
- Tez No: 987954
- Danışmanlar: DR. ÖĞR. ÜYESİ BURCU ÇARKLI YAVUZ
- Tez Türü: Yüksek Lisans
- Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
- Anahtar Kelimeler: Belirtilmemiş.
- Yıl: 2025
- Dil: Türkçe
- Üniversite: Sakarya Üniversitesi
- Enstitü: Fen Bilimleri Enstitüsü
- Ana Bilim Dalı: Bilişim Sistemleri Mühendisliği Ana Bilim Dalı
- Bilim Dalı: Belirtilmemiş.
- Sayfa Sayısı: Belirtilmemiş.
Özet
Bu tez çalışmasında, görüntü tanıma problemlerinde yaygın olarak kullanılan yedi farklı derin öğrenme modeli (ResNet50V2, ResNet152V2, DenseNet121, InceptionV3, Xception, EfficientNetV2S ve MobileNetV2) dört farklı veri seti üzerinde karşılaştırmalı olarak analiz edilmiştir. Tiny ImageNet-200, STL-10, Fashion-MNIST ve TF-Flowers veri setleri kullanılarak modellerin performansları doğruluk, hassasiyet, geri çağırma, F1-skor, ROC-AUC, karmaşıklık matrisi, eğitim süresi, test süresi, toplam parametre sayısı, eğitilebilir parametre sayısı, eğitilemeyen parametre sayısı, inference süresi, RAM ve GPU bellek kullanımı metrikleri ile değerlendirilmiştir. Hiperparametre ayarlamasında Optuna aracı kullanılarak modeller için en uygun öğrenme oranı ve droupout oranı değerleri belirlenmiştir. Deneysel sonuçlar, farklı veri setlerinde performans, maliyet ve kaynak kullanımı yönünden farklı modellerin öne çıktığını göstermektedir. Tiny ImageNet-200 veri setinde Xception modeli 0,6945 doğruluk ve 0,988 ROC-AUC değerleri ile en iyi performansı sergilemiştir. Xception çalışmada kullanılan yedi farklı model arasından Tiny ImageNet-200 gibi karmaşık çok sınıflı görüntü tespiti problemlerde tercih edilebilecek en iyi seçim olmuştur. Modelin yüksek ROC-AUC değeri sınıfları ayırmada yüksek performans göstermektedir. InceptionV3 ve DenseNet121 modelleri de Xception'dan sonra bu veri setinde başarılı sonuçlar elde etmişlerdir. STL-10 veri setinde 0,977 doğruluk ve 0,999 ROC-AUC değerleri ile yine birinci sırada yer alan Xception, bu veri seti için en uygun seçim olmuştur. Düşük çıkarım süresi (0,00089 saniye) gerçek zamanlı uygulamalarda avantaj sağlamaktadır. DenseNet121 Fashion-MNIST veri setinde 0,9394 doğruluk ile en yüksek performansı göstererek, moda öğelerinin özelliklerini içeren bu veri setine benzer veri setlerinde tercih edilebilir bir model olmuştur. TF-Flowers veri setinde ise InceptionV3 ve Xception modelleri öne çıkarak çiçek türlerinin sınıflandırılmasında etkili sonuçlar vermişlerdir. MobileNetV2, tüm veri setlerinde, düşük parametre sayılarında ve RAM kullanımıyla kaynak kısıtlı ortamlar için avantaj sağlamakla birlikte, doğruluk açısından diğer modellerin gerisinde kalmıştır. EfficientNetV2S, performans-verimlilik kalitesi açısından kayda değer sonuçlar elde etmiştir. ResNet modelleri ise parametre sayıları ve uzun eğitim süreleri nedeniyle daha yüksek programlama maliyeti gerektirmiştir. Bu çalışma kapsamında değerlendirilen dört veri seti için belirli öneriler şunlardır: Kaynak sınırlı ortamlarda ve mobil uygulamalarda MobileNetV2 modeli, kabul edilebilir doğruluk oranlarıyla hızlı ve hafif bir çözüm sunmaktadır. Yüksek doğrulukta değişen kritik uygulamalarda ise Xception ve DenseNet121 modelleri tercih edilmelidir. Bu çalışma, kullanılan dört veri seti ve yedi model ile sınırlıdır. Farklı veri koleksiyonları ve yeni nesil modeller ile daha kapsamlı analizlerin yapılması daha genellenebilir sonuçlar verecektir.
Özet (Çeviri)
Scientific progress in image recognition and object classification has accelerated significantly in the past decade, driven predominantly by the rapid advancements in deep learning methods. These achievements have been primarily realized through convolutional neural networks (CNNs), which have redefined the state of the art in computer vision tasks ranging from simple object recognition to complex scene understanding. The sustainability of these advances and their applicability in real-world contexts, however, depend heavily on the careful selection of deep learning architectures that balance predictive performance, computational complexity, and hardware efficiency. Each CNN model—whether ResNet50V2, ResNet152V2, DenseNet121, InceptionV3, Xception, EfficientNetV2S, or MobileNetV2—brings distinct structural innovations, parameter counts, and connectivity patterns. These unique architectural features result in specific advantages and limitations, reinforcing the reality that no single model consistently outperforms others across all datasets, domains, and application requirements. Consequently, systematic benchmarking and comparative analyses have become indispensable tools for both researchers and practitioners seeking to match the most appropriate model to a given problem. The overarching objective of this thesis is to conduct a comprehensive comparative evaluation of seven prominent CNN models across four carefully selected datasets, with the dual aim of contributing to the academic discourse on model benchmarking and offering practical guidance for applied domains such as mobile computing, embedded systems, and large-scale cloud deployments. The study systematically examines ResNet50V2, ResNet152V2, DenseNet121, InceptionV3, Xception, EfficientNetV2S, and MobileNetV2 on the Tiny ImageNet-200, STL-10, Fashion-MNIST, and TF-Flowers datasets, which collectively capture varying levels of complexity, class diversity, and visual representation challenges. By incorporating datasets of differing sizes and domains—from small grayscale fashion items to large-scale natural images—the evaluation seeks to capture the models' generalization abilities under diverse conditions. To evaluate the models comprehensively, this thesis employs a wide range of performance and efficiency metrics. Classical performance indicators such as Accuracy, Precision, Recall, F1-Score, and ROC-AUC are considered essential for quantifying classification success, while additional analysis through confusion matrices enables more nuanced insights into class-level misclassifications. Equally critical, however, are efficiency-oriented metrics that determine the practicality of deploying these models under various hardware constraints. For this reason, training time, testing time, inference time, total and trainable parameter counts, RAM usage, and GPU memory consumption were measured systematically. This dual focus on predictive accuracy and resource efficiency is particularly relevant in contemporary computing contexts, where applications may vary from high-performance computing clusters used in scientific research to severely resource-constrained environments such as smartphones, IoT devices, and autonomous drones. The methodology adopted in this thesis emphasizes fairness and rigor. Hyperparameter optimization was conducted using the Optuna framework, which provides an automated and efficient means of identifying optimal learning rates and dropout probabilities for each model. Transfer learning and fine-tuning strategies were consistently applied, leveraging pre-trained weights to accelerate convergence while enabling adaptation to the target datasets. This methodological consistency ensures that performance differences among the models are attributable primarily to architectural characteristics and not to discrepancies in experimental setup. The experimental results reveal a rich set of insights, confirming that model selection must be context-sensitive. On the Tiny ImageNet-200 dataset, characterized by its complexity and multi-class structure, the Xception model emerged as the strongest performer, achieving an accuracy of 69.45% and a ROC-AUC of 0.988. Its inference time of only 0.00086 seconds further underscores its suitability for real-time classification tasks. DenseNet121 also demonstrated competitive performance, reaching a ROC-AUC of 0.989 with only 7.35 million parameters, highlighting its efficiency in scenarios where computational resources are limited. In contrast, ResNet-based models such as ResNet152V2 and ResNet50V2 required substantially more resources, with parameter counts of 58.9 million and extensive training times, making them less practical for time-sensitive or resource-constrained applications. On the STL-10 dataset, which contains more general object classes, Xception once again dominated, delivering 97.71% accuracy and 0.999 ROC-AUC, paired with a favorable inference time of 0.00089 seconds. This combination of speed and accuracy reinforces Xception's adaptability across diverse domains. MobileNetV2 achieved the fastest inference time of 0.00033 seconds, demonstrating its capacity for real-time deployment, but its lower accuracy of 81.04% indicated trade-offs in predictive reliability, rendering it more appropriate for lightweight applications where speed is prioritized over precision. The Fashion-MNIST dataset provided another dimension to the evaluation by testing models on simple grayscale images of clothing items. Here, DenseNet121 achieved the best balance, with 93.94% accuracy and 0.997 ROC-AUC, confirming its ability to generalize effectively even with relatively simple data. MobileNetV2 also performed well in terms of efficiency, achieving 92.45% accuracy while consuming minimal GPU memory, thereby making it highly suitable for embedded or mobile applications. While ResNet152V2 delivered slightly higher accuracy (94.04%), its significant computational demands again limited its practicality in contexts requiring rapid training or low-latency responses. On the TF-Flowers dataset, which involves distinguishing between various flower species, the Xception model once again led the performance spectrum, achieving 94.82% accuracy and a 0.996 ROC-AUC. DenseNet121 and EfficientNetV2S followed closely, each demonstrating strong generalization with accuracies in the 93–94% range. MobileNetV2, however, performed poorly with only 44.69% accuracy, a result that starkly illustrates the limitations of lightweight models when confronted with complex and nuanced datasets. This outcome underscores the critical need for careful alignment between model architecture and dataset complexity in order to achieve reliable results. These findings collectively highlight several key takeaways. For maximum accuracy applications, particularly those in high-stakes domains such as medical imaging or autonomous vehicle perception, Xception and DenseNet121 stand out as the most reliable choices due to their superior accuracy and ROC-AUC values. Xception further distinguishes itself through its relatively efficient inference times, making it an attractive option for real-time systems. For scenarios where computational and memory resources are constrained, MobileNetV2 represents the best compromise, offering acceptable levels of accuracy while minimizing memory usage and inference latency. From a cost-performance perspective, DenseNet121 consistently demonstrated an optimal balance, delivering high accuracy with low parameter counts and manageable training times. Conversely, the ResNet family, while capable of achieving strong classification results, was hindered by its excessive parameter sizes and resource demands, rendering it impractical for real-time or resource-limited deployment scenarios. Beyond these dataset-specific results, the broader contribution of this thesis lies in establishing a structured framework for evaluating CNN models across both performance and efficiency dimensions. By integrating traditional metrics with resource-oriented measures, the study bridges the gap between academic benchmarking and practical deployment considerations. The findings thus extend their relevance to multiple stakeholders: researchers can use them to deepen theoretical understanding of CNN architectures, practitioners can apply them to optimize deployment in diverse domains, and engineers can leverage them for designing systems that balance accuracy with resource efficiency. In conclusion, this thesis underscores the importance of context-aware model selection in computer vision applications. The comprehensive comparative analysis not only enriches the academic literature by providing systematic benchmarks of widely used CNNs but also delivers actionable insights for real-world implementations. Looking forward, several promising research directions emerge. Future studies may explore the integration of next-generation architectures such as Vision Transformers (ViTs) and multimodal models, which hold potential for further improvements in generalization and cross-domain transferability. Moreover, model compression techniques such as quantization, pruning, and knowledge distillation could be investigated as strategies to reduce computational overhead while preserving accuracy, thereby enabling even broader adoption of advanced CNNs in resource-constrained environments. By building on the findings presented here, the research community can continue advancing towards more efficient, scalable, and versatile computer vision solutions that meet the evolving demands of both academic inquiry and real-world application.
Benzer Tezler
- Derin öğrenme ile cerrahi video anlama
Surgical video understanding with deep learning
ABDISHAKOUR ABDILLAHI AWALE ABDISHAKOUR ABDILLAHI AWALE
Yüksek Lisans
İngilizce
2022
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolGazi ÜniversitesiBilişim Sistemleri Ana Bilim Dalı
DR. ÖĞR. ÜYESİ DUYGU SARIKAYA
- Prediction of COVID 19 disease using chest X-ray images based on deep learning
Derin öğrenmeye dayalı göğüs röntgen görüntüleri kullanarak COVID 19 hastalığının tahmini
ISMAEL ABDULLAH MOHAMMED AL-RAWE
Yüksek Lisans
İngilizce
2024
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolGazi ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
PROF. DR. ADEM TEKEREK
- Facial expression analysis foran online usability evaluation platform
Çevrimiçi kullanılabilirlik değerlendirme platformu için yüz ifadesi analizi
ALİ AZMOUDEH
Yüksek Lisans
İngilizce
2025
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
PROF. DR. HAZIM KEMAL EKENEL
- Medikal görüntü analizinde gürültü saldırılarına karşı derin öğrenme modellerinin performanslarının karşılaştırılması
Benchmarking of deep learning models against adversarial attacks in medical image analysis
GÖKÇE OK
Yüksek Lisans
Türkçe
2024
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolGazi ÜniversitesiBilgi Güvenliği Mühendisliği Ana Bilim Dalı
PROF. DR. MURAT DENER
- Active learning based human in the loop deep object detectionfor scalable data annotation
Ölçeklenebilir veri etiketlenmesi için aktif öğrenme tabanlı insan katılımlı derin nesne tespiti sistemi
ATABERK ARMAN KAYHAN
Yüksek Lisans
İngilizce
2021
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik ÜniversitesiUçak ve Uzay Mühendisliği Ana Bilim Dalı
DOÇ. DR. NAZIM KEMAL ÜRE