Geri Dön

İstanbul sanayi sektörü saatlik elektrik talebi tahmininde bilgi grubu hiyerarşisi ve ablasyon analizi: Xgboost tabanlı bir yaklaşım

Information group hierarchy and ablation analysis in hourly industrial electricity demand forecasting for Istanbul: An xgboost-based approach

  1. Tez No: 1024903
  2. Yazar: REHA BARAN
  3. Danışmanlar: PROF. DR. MUSTAFA BAĞRIYANIK
  4. Tez Türü: Yüksek Lisans
  5. Konular: Elektrik ve Elektronik Mühendisliği, Electrical and Electronics Engineering
  6. Anahtar Kelimeler: Belirtilmemiş.
  7. Yıl: 2026
  8. Dil: Türkçe
  9. Üniversite: İstanbul Teknik Üniversitesi
  10. Enstitü: Lisansüstü Eğitim Enstitüsü
  11. Ana Bilim Dalı: Elektrik Mühendisliği Ana Bilim Dalı
  12. Bilim Dalı: Elektrik Mühendisliği Bilim Dalı
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Elektrik enerjisinin ekonomik ölçekte depolanma imkânının sınırlı olması, üretim ile tüketim arasında eş zamanlı denge gerektirmektedir. Bu zorunluluk, şebeke frekansının nominal aralıkta tutulmasını dengenin sürekliliğine bağlamaktadır. Anlık denge bozulduğunda meydana gelen frekans sapmaları, koruma sistemlerinin devreye girmesine ve kesintilere yol açabilmektedir. Söz konusu zorunluluk, yük tahminini güç sistemi planlama ve operasyonunun en temel girdilerinden biri hâline getirmektedir. Tahmin doğruluğu hem üretim planlamasının ekonomik verimliliğini hem de iletim ve dağıtım altyapısının güvenli işletilmesini doğrudan etkilemektedir. Yük tahmini, enerji sistem planlamasının ilk ve en kritik aşamalarından biri olarak konumlanmaktadır. Geçmiş ve günümüz tüketim örüntülerinin analiz edilmesi ile tahmine etki eden faktörlerin belirlenmesi yoluyla gelecekteki yük seviyelerinin öngörülmesi sürecini ifade etmektedir. Tahmin çalışmaları kapsadığı zaman ufkuna göre çok kısa dönem, kısa dönem, orta dönem ve uzun dönem olarak sınıflandırılmaktadır. Bir saatten kısa ufuklu tahminler çok kısa dönem, bir saat ile bir hafta arasındaki tahminler kısa dönem, bir haftadan bir yıla kadar olanlar orta dönem ve bir yıldan uzun ufuklu tahminler ise uzun dönem yük tahmini olarak nitelendirilmektedir. Bu sınıflandırma içinde kısa dönem yük tahmini operasyonel kararlarda doğrudan rol oynamaktadır. Üretim çizelgesinin oluşturulması, devreye girecek ve devreden çıkacak santrallerin belirlenmesi ile gün öncesi piyasası teklif süreçlerinin desteklenmesi bu kararların başında gelmektedir. Türkiye toplam elektrik tüketiminin önemli bir bölümünü oluşturan sanayi tüketicilerinin yük profili, mesken ve ticari segmentlerden iki temel açıdan ayrışmaktadır. Tüketim büyük ölçüde vardiya çizelgesi, hafta sonu konfigürasyonu ve resmi tatil takvimi gibi operasyonel kararlara bağlı işlemektedir. Klima ve ısıtma yüklerinin ağırlığı mesken segmentine göre düşük kaldığından sıcaklık ile tüketim arasındaki kısa dönem ilişkisi görece zayıftır. Bu iki özellik, sanayi yükü tahmininde değişken seçimi ve özellik mühendisliğinin tasarımını mesken odaklı çalışmalardan ayrıştırmaktadır. Bu çalışma, İstanbul ili sanayi abone grubunun saatlik elektrik talebinin gün öncesi 24 saatlik bir ufukta tahmin edilmesini amaçlamaktadır. Tahmin değişkenleri tekil olarak değil, semantik (anlamsal) olarak ilişkili gruplar hâlinde ele alınmıştır. Yedeklilik sorununun tekil özellik ablasyonunda yarattığı yanıltıcı önemsizlik kanısının önüne geçilmesi için grup bazlı ablasyon çerçevesi kurulmuştur. Çalışmanın merkezinde Sanayi Yükü Katsayısı (ILC) ve altı semantik bilgi grubuna dayalı hiyerarşik özellik çerçevesi yer almaktadır. ILC, sanayi üretim ritminin haftalık örüntüsünü tek bir gösterge altında özetleyen ve tez kapsamında türetilen bir göstergedir. Çalışma 2020–2025 dönemini kapsayan altı yıllık saatlik elektrik tüketim verisi üzerinde yürütülmüştür. Metodolojik akış dört temel aşamada kurgulanmıştır. İlk aşamada üç ekonometrik (OLS, Ridge, LASSO) ve yedi makine öğrenmesi modeli (XGBoost, Gradient Boosting, LightGBM, CatBoost, Extra Trees, Random Forest, AdaBoost) karşılaştırmaya alınmıştır. Modeller, genişleyen pencere ile ardışık doğrulama protokolü altında 2025 yılı boyunca 8.760 saatte sınanmıştır. İkinci aşamada en yüksek doğruluğu sunan baz model üzerinde uzman bilgisi ve literatür temelli değer aralıklarına dayalı manuel hiperparametre ayarlama süreci yürütülmüştür. Üçüncü aşamada altı semantik (anlamsal) bilgi grubu (Bellek, Kültürel, Sektörel, Takvimsel, Meteorolojik, Yapısal) tanımlanmış; bu hiyerarşi üzerinde dört senaryolu ablasyon analizi (S1: Tam Model, S2: Sektörel, S3: Kültürel, S4: Bellek) gerçekleştirilmiştir. Dördüncü aşamada tatil dönemlerinin yıllık tahmin hatası üzerindeki orantısız etkisi nicel olarak ölçülmüş ve hareketli ile sabit tatillerin model davranışı üzerindeki etkileri ayrıştırılarak değerlendirilmiştir. Geliştirilen tahmin hattı ayrıca Python ve MATLAB ortamlarında yeniden çalıştırılarak yazılımdan bağımsızlık testine tabi tutulmuştur. XGBoost on aday model arasında dört doğruluk metriğinin (MAE, RMSE, R², MAPE) tamamında en yüksek performansı sergileyerek baz model olarak benimsenmiştir. Manuel hiperparametre ayarlama süreci sonrasında modelin MAPE değeri %2,71'den %2,36 düzeyine düşürülmüştür. Ablasyon etkilerine göre bilgi grupları Bellek, Kültürel, Sektörel şeklinde sıralanmaktadır. Bellek bilgi grubunun çıkarılması tek başına model performansını kritik biçimde bozarken, diğer grupların çıkarılması daha sınırlı bir etki üretmiştir. Tam modelde en yüksek kazanç (gain) payına sahip bilgi grubu Kültürel iken, ablasyon etkisinde en kritik grup Bellek olarak ortaya çıkmıştır. İki metriğin farklı sorulara yanıt vermesinden kaynaklanan bu uyumsuzluk Kazanç–Ablasyon paradoksu olarak adlandırılmıştır. Tatil bilgi grubu içerisinde hareketli tatil değişkeninin model üzerindeki etkisinin sabit tatil değişkenine göre belirgin biçimde daha yüksek olduğu tespit edilmiştir. Tatil saatleri yıl içinde küçük bir oran oluşturmakla birlikte yıllık tahmin hatasında orantısız bir paya sahiptir. Ablasyon analizi ILC göstergesinin sanayi yük tahmininde Bellek ve Kültürel gruplar tarafından kısmen ikame edilebilen yapısal bir bilgi katmanı sunduğunu ortaya koymuştur. Geliştirilen tahmin hattının iki farklı yazılım ortamında benzer doğruluk seviyeleriyle yeniden üretilebilmesi, tahmin başarısının belirli bir yazılım kütüphanesine değil özellik mühendisliği, hiperparametre seçimi ve doğrulama kurgusunun veriyle uyumundan kaynaklandığını göstermiştir. Çalışmanın bulguları, sanayi sektörünün ayrıştırılmış olarak modellenmesinin ve bilgi gruplarına dayalı hiyerarşik bir çerçevenin tahmin doğruluğunu desteklediğini ortaya koymaktadır.

Özet (Çeviri)

Electricity is an unusual commodity in that it cannot be stored at economically meaningful scales, which compels grid operators to maintain a continuous balance between generation and consumption. Any sustained mismatch between supply and demand causes frequency deviations from the nominal range, triggers protection systems, and may ultimately lead to service interruptions. This physical constraint places load forecasting at the very foundation of power system planning and operations. The accuracy of these forecasts has direct economic and operational consequences: it shapes the efficiency of generation scheduling, the safe operation of transmission and distribution infrastructure, and the financial outcomes of market participants. Improving forecast accuracy is therefore a continuous research priority across utilities, system operators, and energy market actors. Load forecasting is one of the earliest and most critical stages of energy system planning. It refers to the process of estimating future electricity demand by analyzing historical and current consumption patterns and by identifying the factors that influence load behavior. Forecasts are conventionally classified into four categories according to their time horizon. Very short-term forecasts cover horizons shorter than one hour and support real-time balancing. Short-term forecasts span from one hour to one week and underpin generation scheduling and day-ahead market operations. Medium-term forecasts cover periods from one week to one year and inform maintenance planning, while long-term forecasts extend beyond one year and guide capacity expansion decisions. Among these categories, short-term load forecasting plays a particularly active role in everyday operational decision-making. A substantial share of Türkiye's total electricity consumption is attributable to industrial customers, whose load profile differs from residential and commercial segments in two fundamental respects. Consumption is largely driven by operational decisions such as shift schedules, weekend configurations, and the calendar of official holidays. Furthermore, the share of climate-driven loads such as heating and cooling is comparatively low, which weakens the short-term relationship between temperature and consumption. These two features distinguish industrial load forecasting from residential-oriented studies in terms of variable selection and feature engineering design. Istanbul, as one of Türkiye's leading industrial and population centers, provides a particularly relevant setting for examining the hourly demand dynamics of the industrial segment. This thesis aims to forecast the hourly electricity demand of industrial customers in Istanbul over a day-ahead 24-hour horizon. Rather than treating predictor variables in isolation, the study organizes them into semantically related groups and conducts a group-based ablation analysis. This design choice avoids the misleading conclusion of insignificance that single-feature ablation often produces in the presence of variable redundancy. At the core of the methodology lies a hierarchical feature framework built around six semantic information groups: Memory, Cultural, Sectoral, Calendar, Meteorological, and Structural. The framework also introduces the Industrial Load Coefficient (ILC), an original indicator developed within the thesis that captures the weekly rhythm of industrial production through a single normalized value. The study is conducted on six years of hourly electricity consumption data covering the period 2020–2025. The dataset combines load measurements obtained from the Turkish Energy Markets Operator (EPİAŞ) with meteorological observations retrieved from the Open-Meteo platform. The methodological pipeline is structured around four sequential stages. The first stage compares a benchmark pool of ten candidate models. The second stage performs hyperparameter tuning on the best-performing model. The third stage carries out a group-based ablation analysis across four scenarios (S1: Full Model, S2: Sectoral, S3: Cultural, S4: Memory). The fourth stage examines the disproportionate effect of holiday periods on annual forecast error. As an additional validation layer, the entire forecasting pipeline is re-implemented in MATLAB to test its independence from any particular software environment. The benchmark pool consists of three econometric models (Ordinary Least Squares, Ridge, and LASSO regression) and seven machine learning models (XGBoost, Gradient Boosting, LightGBM, CatBoost, Extra Trees, Random Forest, and AdaBoost). The models are evaluated under an expanding-window walk-forward validation protocol that preserves the chronological order of the time series. Each model is retrained on all available historical data up to each prediction day and is then used to forecast the following 24 hours. The procedure produces 365 independent training-testing iterations across 2025, covering 8,760 hourly observations in total. Four accuracy metrics (Mean Absolute Error, Root Mean Squared Error, the coefficient of determination, and Mean Absolute Percentage Error) are reported for each model. The model with the highest accuracy across all four metrics is selected as the baseline, and a manual tuning procedure based on expert knowledge and literature-informed parameter ranges is applied to refine its configuration. The improvement gained through tuning is then verified across monthly slices, the expanding-window learning curve, and the error percentile distribution. The ablation analysis is structured around four scenarios: the full model, the model without the Sectoral (ILC) group, the model without the Cultural (holiday) group, and the model without the Memory (lag) group. In each scenario, only the variables belonging to the targeted information group are removed while all other modeling parameters remain fixed. The shifts in feature importance shares after each ablation are also analyzed to reveal the substitution mechanisms operating between information groups. A dedicated analysis examines the disproportionate contribution of holiday hours to the annual forecast error. Holidays are partitioned into fixed holidays, consisting of the seven official public holidays, and moving holidays, comprising Ramadan and Sacrifice Feasts, which shift across the solar calendar due to their Hijri timing. The two groups are evaluated separately to capture their distinct effects on model behavior. As a final methodological layer, the same forecasting pipeline is reproduced in MATLAB using equivalent algorithms and parameters, and the resulting accuracy metrics are compared with those obtained in Python. XGBoost achieves the highest performance across all four accuracy metrics among the ten candidate models and is selected as the baseline. The manual tuning procedure reduces its MAPE from 2.71% to 2.36%. The ablation analysis reveals a clear hierarchy among information groups, ordered as Memory, Cultural, Sectoral. The Memory group, comprising lag features, emerges as the most critical: its removal sharply degrades model performance, whereas the removal of the other groups produces more limited effects. A notable observation concerns the divergence between feature importance and ablation impact. The Cultural group holds the largest Gain share in the full model, yet the Memory group proves to be the most consequential under ablation. This contrast is termed the Gain–Ablation paradox and reflects the different questions that these two metrics address. The ablation analysis shows that the ILC indicator contributes to forecast accuracy in two complementary ways. While its direct marginal effect on overall accuracy is modest, the indicator plays a decisive role in the substitution mechanism that emerges when the Cultural group is removed from the model. Holiday hours, despite representing only a small fraction of the year, account for a disproportionate share of the annual forecast error, with moving holidays carrying noticeably greater weight than fixed holidays. The forecasting pipeline reproduces consistent accuracy levels across Python and MATLAB environments, confirming that forecast quality stems from feature engineering, hyperparameter selection, and validation design rather than from any specific software library. These findings collectively demonstrate that segmenting the industrial sector separately and adopting a hierarchical information group framework improve short-term load forecasting performance.

Benzer Tezler

  1. Elektrik enerjisi tüketimi, Türkiye değerlendirmesi ve analitik hiyerarşi süreci ile irdelenmesi

    Consumption of electrical energy, Turkish review and study of analitycal hierarchy process

    KEMAL GÖK

    Yüksek Lisans

    Türkçe

    Türkçe

    2016

    Enerjiİstanbul Teknik Üniversitesi

    Enerji Bilimi ve Teknolojileri Ana Bilim Dalı

    PROF. DR. ASİYE BERİL TUĞRUL

  2. Farklı insansız hava araçları ile elde edilen görüntülerin otomatik fotogrametrik yöntemlerle değerlendirilmesi ve doğruluk analizi

    Examination of images obtained from different unmanned air vehicles via automatic photogrammetric methods and accuracy analysis

    DENİZ BİLGE KILINÇOĞLU

    Yüksek Lisans

    Türkçe

    Türkçe

    2016

    Jeodezi ve Fotogrametriİstanbul Teknik Üniversitesi

    Geomatik Mühendisliği Ana Bilim Dalı

    PROF. DR. ELİF SERTEL

  3. Türkiyede tekstil ve konfeksiyon sektörünün durumu ve çıkış stratejileri

    Başlık çevirisi yok

    İBRAHİM ÖZGÜR

    Yüksek Lisans

    Türkçe

    Türkçe

    2006

    Tekstil ve Tekstil MühendisliğiKadir Has Üniversitesi

    Bankacılık Ana Bilim Dalı

    YRD. DOÇ. DR. HASAN EKEN

  4. Analysis of hybrid wind-solar power plant for itu Ayazaga Campus

    İTÜ Ayazağa Yerleşkesi için rüzgar/güneş hibrit güç santralı analizi

    NIMA JAFARZADEH

    Yüksek Lisans

    İngilizce

    İngilizce

    2017

    Enerjiİstanbul Teknik Üniversitesi

    Enerji Bilimi ve Teknolojileri Ana Bilim Dalı

    YRD. DOÇ. DR. BURAK BARUTÇU

  5. Sanayi kuruluşunda çalışanların diyet kalitelerinin iş stresi ve anksiyete üzerine ilişkilerinin incelenmesi

    Investigation of the relationship of diet quality of employees in industrial organization on work stress and anxiety

    RABİA ARAS

    Yüksek Lisans

    Türkçe

    Türkçe

    2022

    Beslenme ve Diyetetikİstanbul Bilgi Üniversitesi

    Beslenme ve Diyetetik Ana Bilim Dalı

    DR. ÖĞR. ÜYESİ BİRSEN DEMİREL