Türkçe e-ticaret yorumlarında transformer tabanlı duygu analizi
Transformer-based sentiment analysis on Turkish e-commerce reviews
- Tez No: 997240
- Danışmanlar: DOÇ. DR. TUNCAY ÖZCAN
- Tez Türü: Yüksek Lisans
- Konular: Endüstri ve Endüstri Mühendisliği, İşletme, Dilbilim, Industrial and Industrial Engineering, Business Administration, Linguistics
- Anahtar Kelimeler: Doğal dil işleme, Metin sınıflandırma, Natural language processing, Text categorization
- Yıl: 2026
- Dil: Türkçe
- Üniversite: İstanbul Teknik Üniversitesi
- Enstitü: Lisansüstü Eğitim Enstitüsü
- Ana Bilim Dalı: İşletme Mühendisliği Ana Bilim Dalı
- Bilim Dalı: İşletme Mühendisliği Bilim Dalı
- Sayfa Sayısı: Belirtilmemiş.
Özet
Bu çalışma, önceden eğitilmiş Transformer mimarilerine dayalı ince ayarlı sınıflandırma modelleri geliştirerek ve Optuna hiperparametre optimizasyonu uygulayarak Türkçe e-ticaret ürün yorumları üzerinde duygu analizi performansını artırmayı amaçlamaktadır. Bu kapsamda sırasıyla veri setinin hazırlanması, model eğitimleri, sonuçların değerlendirilmesi ve önerilen modellerin yayınlanması süreçleri yürütülmüştür. Çalışma kapsamında hazır veri seti kullanılmamış olup; sistematik bir işlem hattı aracılığıyla veri madenciliği, etiketleme ve kapsamlı ön işleme adımları ile güncel ve alan özgü veri seti oluşturulmuştur. Veri kalitesini ve tutarlılığını artırmak amacıyla iş akışına UFAL normalizasyonu, türkçe karakter düzenlemeleri, emoji dönüştürme, url etiket ve email gibi gürültülerin temizlenmesi, harf tekrarlarının azaltılması, noktalama, işaret ve boşluk düzenlemeleri gibi süreçler yürütülmüştür. Ayrıca Türkçe dışı, tekrar eden ve alan dışı yorumlar da filtrelenmiştir. Sonrasında yorum duygu etiketleri Gemma ve LLaMA gibi yerel büyük dil modellerinden yararlanılarak gözden geçirilmiştir. Nihai veri seti 50.160 etiketli yorum ile neticelendiriş olup model eğitimleri sürecinde bağlam daha anlamlı ve kararlı hale getirilmiştir. Hazırlık sürecinin tamamlanmasının ardından son hâline getirilen veri kümesi; BERT, ModernBERT, RoBERTa, TurkishBERTweet, ELECTRA ve DeBERTa gibi güncel Transformer tabanlı temel modellerin hem varsayılan hiperparametrelerle hem de makro-F1 skorunu en üst düzeye çıkarmaya yönelik Optuna tarafından optimize edilen hiperparametre değerleri ile tam ince ayarlı eğitimleri gerçekleştirilmiştir. Deneysel bulgular, Optuna tabanlı hiperparametre optimizasyonu uygulanarak eğitilen modellerin, varsayılan parametrelerle eğitilen modellere kıyasla makro-F1 skorunda yaklaşık %3–%7 oranında iyileşme sağladığını göstermektedir. Test veri setinde en yüksek performans, Optuna ile optimize edilen tam ince ayarlı dbmdz/bert modeli tarafından elde edilmiş olup, model %64,17 makro-F1 ve %91,69 doğruluk değerlerine ulaşmıştır. Buna karşın doğrulama veri setinde en başarılı sonuç, Optuna ile optimize edilen tam ince ayarlı Trendyol/tyroberta modeli tarafından elde edilerek %65,48 makro-F1 ve %92,51 doğruluk skorlarıyla kaydedilmiştir. Transformer bazlı modellerin klasik sınıflandırma algoritmalarına kıyasla daha yüksek Macro F1 sağladığı gözlemlenmiştir. Doğrulama ve test sonuçları arasındaki farkların düşük olması, modellerin genellenebilirliğinin yüksek olduğunu göstermektedir. Genel olarak bu çalışma, modern Transformer mimarilerinin Türkçe e-ticaret bağlamında duygu analizi görevlerinde etkinliğini deneysel olarak ortaya koymakta ve hiperparametre optimizasyonunun model performansı üzerindeki belirleyici etkisini vurgulamaktadır. Yeniden üretilebilirliği sağlamak ve araştırma topluluğuna açık bir kaynak sunmak amacıyla, geliştirilen modeller Hugging Face platformunda kamuya açık şekilde yayımlanmıştır.
Özet (Çeviri)
With the rapid expansion of digitalization, e-commerce platforms have become one of the most influential sources of information shaping consumers' purchasing behaviors. User-generated product reviews not only serve as a reference point for potential buyers but also provide critical insights for manufacturers and sellers in terms of measuring customer satisfaction, managing product development processes, and gaining a competitive advantage. In this context, the automatic processing of large-scale and unstructured textual data to extract meaningful insights has brought sentiment analysis—widely studied within the field of Natural Language Processing (NLP)—to the forefront. However, e-commerce reviews inherently exhibit a noisy structure due to spelling errors, grammatical inconsistencies, slang expressions, and the extensive use of emojis, which significantly limits the effectiveness of traditional classification algorithms. Turkish, with its agglutinative morphological structure and rich word-formation capacity, presents unique challenges for sentiment analysis compared to widely studied languages such as English. The ability of word stems to remain unchanged while acquiring numerous suffixes substantially increases vocabulary size and leads to data sparsity issues. Although the literature indicates a growing number of studies on Turkish sentiment analysis, a clear research gap remains—particularly regarding the systematic and experimental investigation of handling class imbalance in e-commerce-specific datasets and the impact of hyperparameter optimization on model performance. Motivated by this gap, this thesis aims to enhance the effectiveness of modern Transformer architectures on Turkish e-commerce data and to empirically demonstrate the role of hyperparameter optimization (HPO) in this process. One of the most critical factors determining the success of sentiment analysis is the quality of the data. Accordingly, instead of relying on an existing dataset, the dataset used in this study was constructed from scratch to accurately reflect real-world scenarios. The data mining process was conducted on the“Sunglasses”category of a JavaScript-based e-commerce platform using Python-based Selenium and Beautiful Soup libraries. As a result of crawling 11,014 products across 242 brands, a total of 159,552 raw user reviews were collected. The raw data was found to be highly imbalanced (predominantly positive) and noisy due to the nature of e-commerce reviews, which often contain spelling errors, emojis, and repeated characters. To prepare the dataset for model training, systematic labeling, preprocessing, and filtering strategies were applied. To mitigate class imbalance during the labeling process, the majority of neutral and negative reviews and a subset of positive reviews were selected, resulting in a labeled dataset of 60,297 samples. A comprehensive preprocessing pipeline was implemented to improve data quality and ensure suitability for model training. In particular, the UFAL (MultiLexNorm) model—winner of the W-NUT 2021 shared task and based on ByT5 character-level normalization—was employed to correct frequent spelling errors commonly observed in e-commerce product reviews. Due to limitations in the pretraining data of the UFAL model, Turkish character corrections were further supported using the Zemberek Deasciify library. Emojis, which play a crucial role in sentiment expression, were not removed but instead converted into textual representations to preserve their semantic meaning. Additional preprocessing steps included the removal of URLs, tags, and email addresses; reduction of repeated characters; and normalization of punctuation, symbols, and whitespace. Furthermore, non-Turkish, duplicate, and domain-irrelevant reviews were filtered to further refine the dataset. In addition to manual labeling, predictions were generated using local large language models (LLMs), namely Gemma3:4b and LLaMA3.2:3b, on the fully preprocessed dataset to minimize human error and enhance consistency. Manual labels, LLM predictions, and user ratings were jointly evaluated, resulting in the identification of 1,101 inconsistently labeled reviews. These samples were manually reviewed, and final sentiment labels (Positive, Negative, Neutral) were assigned. Following labeling, preprocessing, and filtering, the final dataset consisted of 50,160 labeled reviews, yielding a more stable and semantically coherent context for model training. The dataset was split into training (68%), validation (12%), and test (20%) sets, ensuring that class distributions were preserved across all subsets through stratified splitting. After completing the dataset preparation phase, six Transformer-based base models—BERT, ModernBERT, RoBERTa, TurkishBERTweet, ELECTRA, and DeBERTa—were selected based on criteria such as architectural design, training strategies, case sensitivity (cased/uncased), Turkish language support, vocabulary size, tokenizer characteristics, model recency, popularity, and compatibility with social media and e-commerce language. Additionally, classical classification algorithms, including Complement Naive Bayes and k-Nearest Neighbors, were incorporated to provide a comprehensive performance comparison framework. One of the key methodological strengths of this study lies in performing fine-tuning not only with default hyperparameter values but also through a systematic hyperparameter optimization process. In Transformer-based models, hyperparameters such as learning rate, batch size, warmup ratio, and weight decay have a substantial impact on performance. Unlike grid search or random search methods, Bayesian Optimization offers more efficient exploration by learning from previous evaluations, reducing computational cost, and incorporating early stopping strategies. Therefore, the Bayesian optimization-based Optuna framework was employed. For each model, 100 trials were conducted to identify the optimal hyperparameter combinations that maximize the Macro-F1 score, which was selected as the primary objective metric due to class imbalance. Final training was completed using these optimized hyperparameter values. All experiments were conducted locally on an NVIDIA GeForce RTX 3070 Ti GPU, and results were evaluated using both overall and class-level metrics. Experimental results demonstrate that models trained with Optuna-based hyperparameter optimization achieve approximately 3–7% improvements in Macro-F1 scores compared to models trained with default hyperparameter values. The highest performance on the test dataset was achieved by the Optuna-optimized, fully fine-tuned Dbmdz/bert-base-turkish-128k-uncased model, attaining a Macro-F1 score of 64.17% and an accuracy of 91.69%. Conversely, the best performance on the validation dataset was obtained by the Optuna-optimized, fully fine-tuned Trendyol/tyroberta model, with a Macro-F1 score of 65.48% and an accuracy of 92.51%. The success of these two models (BERT- and RoBERTa-based) clearly highlights the advantages of large vocabularies in morphologically rich languages such as Turkish (BERT-128k) and domain-specific pretraining on e-commerce data (TyRoBERTa). Social media-oriented TurkishBERTweet and multilingual mDeBERTa models also achieved competitive results, albeit slightly below the leading models. Overall, Transformer-based models consistently outperformed classical classification algorithms in terms of Macro-F1 performance. The minimal discrepancy between validation and test results further indicates strong model generalizability. Class-level analyses reveal that the models performed exceptionally well in identifying Positive (F1 > 95%) and Negative (F1 > 80%) classes. Notably, the high accuracy in detecting negative reviews is particularly valuable for organizations aiming to capture customer complaints effectively. However, identifying the Neutral class remained a challenging task across all models (F1 < 20%). This difficulty can primarily be attributed to class imbalance and the inherently high semantic ambiguity of neutral expressions. In summary, this study empirically demonstrates the effectiveness of modern Transformer architectures in Turkish e-commerce sentiment analysis tasks and underscores the decisive role of hyperparameter optimization in improving model performance. The proposed models offer direct practical value for decision-making processes such as product management, customer experience analysis, and product assortment planning. Automating the consistent analysis of tens of thousands of customer reviews—previously requiring manual inspection—provides organizations with significant time and cost advantages while enabling faster and more accurate identification of customer expectations. Moreover, jointly analyzing sentiment distributions alongside product attributes, price levels, and user profiles facilitates data-driven decision-making in marketing strategies, product improvement initiatives, and portfolio optimization. Consequently, sentiment analysis transcends its traditional role as a customer satisfaction measurement tool and emerges as a strategic decision-support mechanism for product development, churn prediction, marketing strategy formulation, and competitive advantage creation. To ensure reproducibility and provide an open resource for the research community, the proposed models have been publicly released on the Hugging Face platform. This initiative enables both the Turkish NLP community and industry practitioners to directly utilize or further develop the pretrained models in their own projects. Future work includes comparing different hyperparameter optimization techniques during model fine-tuning and addressing class imbalance by augmenting the dataset through synthetic data generation and advanced data mining approaches. Additionally, leveraging the developed models to perform sentiment inference on unlabeled datasets opens avenues for building recommendation systems based on customer sentiment.
Benzer Tezler
- E-ticaret sektöründe derin öğrenme tabanlı duygu analizi: hiperparametre optimizasyonu ve performans karşılaştırması
Deep learning-based sentiment analysis in the E-commerce sector: hyperparameter optimization and performance comparison
MİRAÇ ÖZTÜRK
Yüksek Lisans
Türkçe
2026
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolÜsküdar ÜniversitesiYapay Zeka Ana Bilim Dalı
PROF. DR. SERHAT ÖZEKES
- Türkçe E-ticaret yorumlarının çok etiketli analizi için derin öğrenme modellerinin uygulanması
Applying deep learning models for multi-label analysis of Turkish E-commerce comments
ABDULKADİR ŞEN
Yüksek Lisans
Türkçe
2025
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolKonya Teknik ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
DR. ÖĞR. ÜYESİ FATMA ZEHRA SOLAK
- Amazon müşteri yorumlarının duygu analizi yöntemleriyle değerlendirilmesi
Evaluating Amazon customer reviews through sentiment analysis techniques
SABUHI YUSIFOV
Yüksek Lisans
Türkçe
2024
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolOstim Teknik ÜniversitesiYazılım Mühendisliği Ana Bilim Dalı
PROF. DR. ALİ SEBETCİ
- Derin öğrenme tabanlı ve transformer tabanlı dil modelleri ile e-ticaret yorumlarının sınıflandırılması
E-commerce reviews classification with deep learning-based and transformer-based language models
BURCU GEL
Yüksek Lisans
Türkçe
2026
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolKto Karatay ÜniversitesiElektrik ve Elektronik Mühendisliği Ana Bilim Dalı
DOÇ. DR. ALİ ÖZTÜRK
- E-ticaret sistemlerinde yapılan ürün yorumlarının metin madenciliği uygulaması ile incelenmesi
Analysis of product comments are examined in e-commerce systems with text mining application
GÖKAY YILMAZ
Yüksek Lisans
Türkçe
2021
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrolİstanbul Aydın ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
PROF. DR. ZAFER ASLAN