Otonom araçların şerit algılama sistemlerinde YOLO mimarilerinin (v9, v11, v12, v26) semantik segmentasyon performanslarının karşılaştırmalı değerlendirilmesi
Comparative evaluation of the semantic segmentation performance of YOLO architectures (v9, v11, v12, v26) in lane detection systems of autonomous vehicles
- Tez No: 1017139
- Danışmanlar: PROF. DR. İSMAİL HAKKI CEDİMOĞLU
- Tez Türü: Yüksek Lisans
- Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
- Anahtar Kelimeler: Otonom Sürüş, Şerit Algılama, Derin Öğrenme, YOLO Mimarileri, Görüntü Segmentasyonu, Semantik Segmentasyon, Autonomous Driving, Lane Detection, Deep Learning, YOLO Architectures, Image Segmentation, Semantic Segmentation
- Yıl: 2026
- Dil: Türkçe
- Üniversite: Sakarya Üniversitesi
- Enstitü: Fen Bilimleri Enstitüsü
- Ana Bilim Dalı: Bilişim Sistemleri Mühendisliği Ana Bilim Dalı
- Bilim Dalı: Belirtilmemiş.
- Sayfa Sayısı: Belirtilmemiş.
Özet
Otonom araç teknolojilerinin gelişimi, modern ulaşım sistemlerinde güvenliği maksimize etme ve trafik akışını optimize etme yolunda devrimsel bir rol oynamaktadır. Bu sistemlerin çevresel algılama katmanındaki en kritik bileşenlerden biri olan şerit algılama; aracın yol üzerindeki yörüngesini belirlemesi, şeritte kalma asistanı (LKA) ve güvenli şerit değiştirme manevraları için temel navigasyon verisini sağlar. Araçların çevresini algılamak için güvenli ve verimli teknolojiler olan şerit takip ve otonom sistemler, yoldaki şeritlerin doğru tespiti ve engellerin tespiti gibi önemli görevleri üstlenerek otonom sürüş için güvenli bir temel sağlar. Bu tez çalışmasında, otonom sürüş sistemlerinde güvenli navigasyonun temel taşlarından biri olan şerit algılama problemini çözmek amacıyla, derin öğrenme tabanlı güncel segmentasyon mimarilerinin performansları kapsamlı bir şekilde incelenmiştir. Otonom araçlarda güvenli navigasyon için şeritlerin kutu içine alınmasının yanı sıra piksel bazında hatasız algılanması da gerekmektedir. Literatürde genellikle tek bir modele odaklanılırken, bu tez çalışmasında hızla gelişen YOLO ailesi v9'dan en güncel v26'ya kadar eş zamanlı olarak incelenmiştir. Tez çalışmasının amacı mAP değerlerinin ölçülmesinin yanı sıra, donanım kısıtlamalarına göre en iyi 'performans/maliyet' dengesini sunan modelin belirlenmesidir. Literatürde YOLO serisinin bu kadar geniş bir yelpazede (v9'dan v26'ya) karşılaştırıldığı çalışma sayısı oldukça sınırlıdır. Araştırma kapsamında, en yeni nesil nesne tespiti ve segmentasyon modelleri olan YOLOv9, YOLOv11, YOLOv12 ve YOLOv26 mimarileri, özgün bir yol segmentasyonu veri seti üzerinde uçtan uca eğitilmiş ve test edilmiştir. Otonom araçların şerit içinde daha istikrarlı ve güvenli bir şekilde hareket etmesini sağlamak için görüntü işleme tabanlı bir şerit takip sistemi geliştirilmiştir. Sistemde, KITTI Road/Lane Detection Dataset (224 X 224) görüntüleri ile yol şeritlerini tespit etmek için segmentasyon teknikleri kullanılmıştır. Bu sistemde, road ve background olmak üzere iki sınıf kullanılmıştır. Literatürde nesne tespiti ve segmentasyon üzerine yapılan çalışmalar genellikle tek bir modelin başarısına odaklanırken, hızla gelişen YOLO serisinin (v9'dan v26'ya kadar) eş zamanlı ve karşılaştırmalı analizi konusundaki veriler sınırlıdır. Mevcut çalışmaların çoğu sadece doğruluk (mAP) değerlerine odaklanırken, bu tezde sunulan görsel tahminler (val_batch_pred) ve korelogram grafikleri, modellerin otonom şerit algılamada maske kenar pikselleri bazındaki“pürüzsüzlük”ve“güven skorları”arasındaki ilişki detaylandırılmıştır. Özellikle YOLOv26s-seg modelinin çok daha az eğitim döngüsünde (epoch) yüksek başarıya ulaşması, gerçek zamanlı sistemler için“eğitim maliyeti/performans”dengesi açısından yeni bir referans noktası sunduğunu kanıtlamıştır. Otonom sistemlerin donanım kısıtlamalarına (FPS ihtiyacı) ve güvenlik gereksinimlerine (yüksek hassasiyet) göre hangi YOLO versiyonunun hangi senaryoda en iyi sonuç verdiği bilimsel verilerle ortaya koyulmuştur. Road ve background ayrımı sadece bir kutu tespiti olarak değil, pikseller bazında en az hata ile otonom aracın yol sınırlarının santimetre hassasiyetinde algılanması sağlanmıştır. Veri seti analizleri, road ve background sınıfları arasında sağlanan sayısal dengenin modellerin genelleme yeteneği üzerinde kritik bir rol oynadığını göstermiştir. Deneysel sonuçlar ve metrik analizleri incelendiğinde; YOLOv9e-seg modelinin mAP50 ve hassas maske tespiti (mAP50-95) değerlerinde en kararlı ve yüksek başarıyı sergilediği görülürken, YOLOv26s-seg modelinin çok daha az eğitim döngüsünde (epoch) benzer doğruluk seviyelerine ulaşarak yüksek yakınsama hızı sunduğu tespit edilmiştir Karmaşıklık matrisleri ve görsel tahmin çıktıları üzerinden yapılan değerlendirmeler sonucunda, modellerin %95'in üzerinde bir güven skoruyla yol sınırlarını pikseller bazında başarıyla ayırt edebildiği kanıtlanmıştır. Çalışma, otonom araçlarda kullanılacak donanımın işlem kapasitesine göre; maksimum doğruluk için YOLOv9 ve YOLOv12, gerçek zamanlı verimlilik için ise YOLOv26 modellerinin tercih edilebileceği bu çalışma ile gösterilmiştir. Kutu (Box) ve maske (Mask) olmak üzere iki çıktı türü performans kriteri değerleri elde edilmiştir. Box çıktısı için tüm modeller göz önüne alındığında, en yüksek kesinlik değeri 0,96 ile YOLOv9c-seg, en yüksek duyarlılık değeri 0,93 ile YOLOv9e-seg, en yüksek F1-Skor değeri 0,94 ile YOLOv9e-seg, en yüksek [email protected] değeri 0,97 ile YOLOv9e-seg, en yüksek [email protected]:0.95 değeri 0,86 ile YOLOv11x-seg modellerine aittir. Mask çıktısı için tüm modeller göz önüne alındığında, en yüksek kesinlik değeri 0,97 ile YOLO26s-seg, en yüksek duyarlılık değeri 0,94 ile YOLOv9e-seg, en yüksek F1-Skor değeri 0,94 ile YOLOv9e-seg, en yüksek [email protected] değeri 0,96 ile YOLOv9e-seg, en yüksek [email protected]:0.95 değeri 0,79 ile YOLOv11x-seg modellerine aittir. Bu değerler incelendiğinde, box ve mask için en doğru değerleri veren YOLOv9e-seg modeli olduğu anlaşılmıştır. YOLOv9e-seg, duyarlılık, F1-Skor, [email protected] değerleri box ve mask değerlerinde doğruluk oranı en yüksektir. [email protected]:0.95 değerleri açısından ise box ve mask ikilisinde en iyi doğruluk oranının YOLOv11x-seg modeline ait olduğu ortaya koyulmuştur. Araştırma kapsamında YOLO11s-seg ve YOLO26s-seg olmak üzere 2 modeli cross validation yöntemi ile eğittik. Yol segmentasyonu performansı açısından karşılaştırılan iki güncel derin öğrenme mimarisinden YOLO11s-seg, hem mimari verimlilik hem de segmentasyon hassasiyeti noktalarında YOLO26s-seg modeline göre daha üstün sonuçlar sergilemiştir. YOLO11s-seg mimarisi, daha az katman (114 katman) ve daha düşük hesaplama yüküyle (32,8 GFLOPs) çalışmasına rağmen, 0,86 Mask mAP50-95 değerine ulaşarak nesne sınırlarını belirlemede yüksek bir başarı göstermiştir. Buna karşın YOLO26s-seg modeli, 139 katmanlı daha derin bir ağ yapısına ve 34,1 GFLOPs işlem karmaşıklığına sahip olup, 0,814 Mask mAP50-95 skoru ile güçlü ancak YOLO11s-seg'in gerisinde kalan bir performans ortaya koymuştur. Elde edilen bulgular, otonom sürüş ve yol analizi gibi gerçek zamanlı uygulama gereksinimleri için YOLO11s-seg modelinin, daha düşük donanım maliyeti ile daha keskin segmentasyon maskeleri üretebilen, optimize bir çözüm sunduğunu doğrulamaktadır. Araştırma bulguları ile YOLO serisinin düşük gecikme süresi gerektiren kritik anlık tepki senaryolarındaki mutlak üstünlüğü kanıtlanmıştır. Elde edilen sonuçlar ile otonom araçlarda kullanılacak algılama sistemlerinin tasarımında hangi model hiyerarşisinin hangi çevresel koşullar için optimize edilmesi gerektiğine dair kapsamlı bir mühendislik rehberi sunulmuştur. Çalışma, gelecekteki otonom sürüş mimarileri için yüksek hızlı şerit algılama ile derin anlamsal segmentasyonu birleştiren hibrit yaklaşımların potansiyeline dikkat çekerek literatüre özgün bir katkı sağlamıştır.
Özet (Çeviri)
The development of autonomous vehicle technologies is playing a revolutionary role in maximizing safety and optimizing traffic flow in modern transportation systems. Lane detection, one of the most critical components in the environmental perception layer of these systems, provides essential navigation data for determining the vehicle's trajectory on the road, lane keeping assist (LKA), and safe lane change maneuvers. Lane keeping and autonomous systems, which are safe and efficient technologies for perceiving the surroundings of vehicles, provide a safe foundation for autonomous driving by undertaking important tasks such as accurate lane detection and obstacle detection on the road. In this thesis, the performance of current deep learning-based segmentation architectures is comprehensively examined to solve the lane detection problem, one of the cornerstones of safe navigation in autonomous driving systems. Safe navigation in autonomous vehicles requires not only the enclosure of lanes but also pixel-by-pixel error-free detection. While the literature generally focuses on a single model, this thesis examines the rapidly evolving YOLO family from v9 to the most recent v26 simultaneously. The aim of this thesis is to measure mAP values and to determine the model that offers the best 'performance/cost' balance according to hardware constraints. The number of studies in the literature comparing the YOLO series across such a wide range (from v9 to v26) is quite limited. In this research, the latest generation object detection and segmentation models, YOLOv9, YOLOv11, YOLOv12, and YOLOv26 architectures, were trained and tested end-to-end on a unique road segmentation dataset. An image processing-based lane tracking system was developed to enable autonomous vehicles to move more stably and safely within their lanes. The system uses segmentation techniques to detect road lanes using KITTI Road/Lane Detection Dataset (224 x 224) images. Two classes, road and background, were used in this system. While studies on object detection and segmentation in the literature generally focus on the success of a single model, data on the simultaneous and comparative analysis of the rapidly evolving YOLO series (from v9 to v26) is limited. Most existing studies focus only on accuracy (mAP) values, whereas the visual predictions (val_batch_pred) and correlogram graphs presented in this thesis detail the relationship between“smoothness”and“confidence scores”of the models in autonomous lane detection based on mask edge pixels. In particular, the high success rate of the YOLOv26s-seg model in significantly fewer training cycles (epochs) proves to offer a new benchmark in terms of“training cost/performance”balance for real-time systems. Scientific data reveals which YOLO version performs best in which scenario, considering the hardware constraints (FPS requirement) and safety requirements (high accuracy) of autonomous systems. The distinction between road and background is not merely a box detection; it allows the autonomous vehicle to perceive road boundaries with centimeter precision, at the pixel level, with minimal error. Dataset analyses have shown that the numerical balance achieved between road and background classes plays a critical role in the generalization ability of the models. Examination of experimental results and metric analyses reveals that the YOLOv9e-seg model exhibits the most stable and high success in mAP50 and precise mask detection (mAP50-95), while the YOLOv26s-seg model achieves similar accuracy levels in significantly fewer training cycles (epochs), offering a high convergence speed. Evaluations based on complexity matrices and visual prediction outputs demonstrate that the models successfully distinguish road boundaries at the pixel level with a confidence score above 95%. This study shows that, depending on the processing capacity of the hardware to be used in autonomous vehicles, YOLOv9 and YOLOv12 models can be preferred for maximum accuracy, while YOLOv26 models can be preferred for real-time efficiency. Two output types, Box and Mask, were used to obtain performance criterion values. For the Box output, considering all models, the highest accuracy value was 0.96 with YOLOv9c-seg, the highest sensitivity value was 0.93 with YOLOv9e-seg, the highest F1-Score value was 0.94 with YOLOv9e-seg, the highest [email protected] value was 0.97 with YOLOv9e-seg, and the highest [email protected]:0.95 value was 0.86 with YOLOv11x-seg. When all models are considered for the mask output, the highest accuracy value is 0.97 for YOLO26s-seg, the highest sensitivity value is 0.94 for YOLOv9e-seg, the highest F1-Score value is 0.94 for YOLOv9e-seg, the highest [email protected] value is 0.96 for YOLOv9e-seg, and the highest [email protected]:0.95 value is 0.79 for YOLOv11x-seg. Examination of these values reveals that the YOLOv9e-seg model provides the most accurate values for both box and mask. YOLOv9e-seg has the highest accuracy rate for both box and mask values in terms of sensitivity, F1-Score, and [email protected]. In terms of [email protected]:0.95, the best accuracy rate for both box and mask pairs belongs to the YOLOv11x-seg model. In this study, the performances of YOLO26s-seg and the next-generation YOLO11s-seg architectures were comprehensively analyzed within the scope of road segmentation, a critical task for autonomous driving systems. To evaluate the generalization capacity of the models, a 5-fold cross-validation strategy was implemented, and the data obtained during the first phase (Fold 0) were evaluated according to academic standards. The YOLO26s-seg model, consisting of 139 layers and 10,366,114 parameters, demonstrated significant proficiency in the 'Road' class with a 0.973 mAP50 and 0.814 Mask mAP50-95 accuracy rate. Conversely, the YOLO11s-seg architecture, which features a more compact network structure of 114 layers, 10,067,590 parameters, and a computational complexity of 32.8 GFLOPs, exhibited a notable superiority in both memory efficiency and inference time. Comparative performance results indicated that the YOLO11s-seg model achieved a Mask mAP50-95 score of 0.86, distinguishing object boundaries with higher precision than the previous generation, and minimized the false positive rate with a high Box Precision value of 0.979. Consequently, the fact that the YOLO11s-seg architecture provides higher segmentation performance despite a lower computational load proves that the model offers an optimized solution for autonomous systems in terms of both operational speed and accuracy. The study makes a unique contribution to the literature by highlighting the potential of hybrid approaches combining high-speed lane detection with deep semantic segmentation for future autonomous driving architectures.
Benzer Tezler
- Generating a high-definition map (HD MAP) via yolo (you only look once) deep learning-based object detection model
Yolo (you only look once) derin öğrenme tabanli nesne tespit modeli ile yüksek tanimli hari̇ta (HD MAP) oluşturma
YASİN MEMİ
Yüksek Lisans
İngilizce
2025
Mühendislik Bilimleriİstanbul Teknik ÜniversitesiGeomatik Mühendisliği Ana Bilim Dalı
PROF. DR. HANDE DEMİREL
- 3D nokta bulutu verileri kullanılarak otonom sürüş için nesne algılama yöntemi ile karayolu envanterlerinin tespit edilmesi
Determination of highway inventories with object detection method for autonomous driving using 3D point cloud data
HİLAL GEZGİN
Doktora
Türkçe
2025
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik ÜniversitesiBilişim Uygulamaları Ana Bilim Dalı
PROF. DR. REHA METİN ALKAN
- Reducing in-vehicle communication overload and enhancing efficiency in autonomous and electrical vehicles
Otonom ve elektrikli araçlarda araç içi iletişim yükünü azaltma ve etkinliğini artırma
YUNUS KAĞAN ÖZDEMİR
Yüksek Lisans
İngilizce
2024
Otomotiv Mühendisliğiİstanbul Teknik ÜniversitesiElektrik Mühendisliği Ana Bilim Dalı
PROF. DR. AHMET CANSIZ
- Model reference adaptive controller design with augmented error method for lane tracking
Serit takibi kontrolü için artıtılmış hata yöntemi ile model referans uyarlanabilir kontrolör tasarımı
MEHMET NURİ DİYİCİ
Yüksek Lisans
İngilizce
2023
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik ÜniversitesiMekatronik Mühendisliği Ana Bilim Dalı
DOÇ. DR. YAPRAK YALÇIN
- Şerit takip desteği sistemi için fonksiyonel emniyet analizi
Functional safety analysis for lane keeping assistance system
EMİR KUDUN
Yüksek Lisans
Türkçe
2023
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Teknik ÜniversitesiKontrol ve Otomasyon Mühendisliği Ana Bilim Dalı
DOÇ. DR. İLKER ÜSTOĞLU