Türkçe varlık ismi bağlama
Turkish named entity linking
- Tez No: 1008615
- Danışmanlar: DOÇ. DR. AHMET CÜNEYD TANTUĞ
- Tez Türü: Yüksek Lisans
- Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
- Anahtar Kelimeler: Doğal dil işleme, Varlık isimleri, Veri işleme, Veri setleri, Yapay zeka, Natural language processing, Named entities, Data processing, Data sets, Artificial intelligence
- Yıl: 2026
- Dil: Türkçe
- Üniversite: İstanbul Teknik Üniversitesi
- Enstitü: Lisansüstü Eğitim Enstitüsü
- Ana Bilim Dalı: Bilgisayar Mühendisliği Ana Bilim Dalı
- Bilim Dalı: Bilgisayar Mühendisliği Bilim Dalı
- Sayfa Sayısı: Belirtilmemiş.
Özet
Yapısal olmayan dijital metinlerin hacmindeki hızlı artış, bu verilerin anlamlandırılmasını Doğal Dil İşleme alanının temel araştırma problemlerinden biri haline getirmiştir. Bu tez çalışmasında, bilgi çıkarımı sürecinin önemli bir aşaması olan Varlık İsmi Bağlama problemi Türkçe özelinde incelenmiştir. Çalışma kapsamında öncelikle Türkçe Varlık İsmi Bağlama literatüründeki nitelikli veri kümesi eksikliğini gidermek amacıyla; Vikipedi ve Vikiveri tabanlı, zengin dilbilimsel özellikler içeren geniş ölçekli“TurkLink”veri kümesi oluşturularak açık kaynak olarak araştırmacıların erişimine sunulmuştur. Veri kümesinin kalitesini artırmak ve model eğitimine uygun hale getirmek üzere iki aşamalı bir filtreleme mekanizması geliştirilmiştir. Kural tabanlı filtreleme aşamasında, Vikiveri'nin hiyerarşik bilgi grafiği üzerindeki ontolojik ilişkiler kullanılarak hedeflenen on bir varlık sınıfı ile ön eşleştirmeler gerçekleştirilmiştir. Kural tabanlı yaklaşımın neden olabileceği bağlamsal belirsizlikleri gidermek amacıyla uygulanan model tabanlı filtreleme aşamasında ise, Büyük Dil Modelleri tabanlı bir topluluk yapısı kullanılarak kelimelerin cümle içindeki bağlamı analiz edilmiş ve çoğunluk oylaması yöntemiyle nihai etiketler belirlenmiştir. Modelleme aşamasında, literatürdeki popüler ReFinED Varlık İsmi Bağlama mimarisi BERTurk ön-eğitimli dil modeliyle entegre edilerek Türkçeye uyarlanmıştır. Gerçekleştirilen deneylerde, varlık sınırlarının önceden verildiği Anlam Belirsizliği Giderme senaryolarında sistemin hem TurkLink hem de Mewsli-9 test kümelerinde %90'ın üzerinde F1-skoru ile yüksek bir başarım gösterdiği gözlemlenmiştir. Buna karşın, varlık sınırlarının model tarafından tahmin edildiği uçtan uca Varlık İsmi Bağlama senaryolarında, Varlık İfadesi Tespiti aşamasının genel sistem performansını sınırlandıran bir darboğaz oluşturduğu görülmüştür. Bu problemi çözmek amacıyla Varlık İsmi Bağlama görevi; Varlık İfadesi Tespiti ve Anlam Belirsizliği Giderme görevleri olarak ayrıştırılmış, harici bir BERTurk tabanlı Varlık İfadesi Tespiti modelinin sürece dahil edildiği ardışık bir mimari oluşturulmuştur. Gerçekleştirilen karşılaştırmalı analizler ve harici çapraz alan doğrulama testleri sonucunda; uçtan uca çalışan bütünleşik modellerin TurkLink test kümesinde 0,3478, haber metinlerinden oluşan Mewsli-9 test kümesinde ise 0,2927 F1-skorunda kaldığı görülmüştür. Önerilen ardışık mimarinin ise TurkLink'te 0,5317, Mewsli-9 kümesinde ise 0,4773 F1-skoruna ulaşarak modelin farklı metin türlerindeki genellenebilirlik kapasitesini önemli ölçüde artırdığı gözlemlenmiştir. Diğer bir taraftan, geliştirilen Türkçe odaklı ardışık model, Mewsli-9 testlerinde literatürdeki en iyi çok dilli modelleri açık ara geride bırakmıştır. Çok dilli modeller Türkçenin yapısında zorlanıp çok fazla hatalı tahmin yaparak en fazla 0,2960 F1-skorunda kalırken; çalışma kapsamında sunulan sistem 0,4773 F1-skoruna ulaşmış ve en güçlü rakibine 0,1813 puan farkla büyük bir fark atmıştır. Sonuç olarak, Türkçe Varlık İsmi Bağlama uygulamalarında görevlerin ardışık bir mimariyle gerçekleştirilmesinin, uçtan uca yaklaşımlara kıyasla daha kararlı, esnek ve yüksek başarımlı bir çözüm sunduğu sonucuna ulaşılmıştır.
Özet (Çeviri)
With the rapid spread of information technology and digital platforms, the global volume of data has increased massively. Most of this data consists of unstructured free text, such as social media posts, news articles, and blogs. Information extraction, the process of converting this raw data into a machine-readable, meaningful, and structured format, is a fundamental research area in Natural Language Processing. This thesis focuses on Entity Linking in the Turkish language, which is one of the most critical and complex stages of information extraction. Entity Linking is the process of detecting the boundaries of entity names (people, organizations, locations, etc.) in a text, resolving their meaning through contextual analysis, and matching them with exact references in massive knowledge bases like Wikidata. For example, resolving whether the word“Fatih”refers to a historical figure (Mehmed II) or a district in Istanbul is a core problem that Entity Linking systems aim to solve. A review of the literature shows that high-performance Entity Linking systems are mostly built for English, while comprehensive, high-quality, and deeply labeled resources tailored to the structure of Turkish are extremely limited. The main goal of this thesis is to address this gap by creating a large-scale Turkish dataset and developing innovative, high-performance Entity Linking models suited to Turkish language. To achieve these goals, the first stage of the thesis introduces“TurkLink,”a rich dataset based on Wikipedia and Wikidata, which has been open-sourced for researchers to contribute to the Turkish Natural Language Processing literature. The dataset was created using the April 2024 Turkish Wikipedia dumps. Articles were cleaned of non-text noise like infoboxes, tables, and templates, and converted into plain text. Wikipedia hyperlinks formed the basis for labeling, as they provide a natural bridge between the text and the entity. The target knowledge base was Wikidata, which represents entities with language-independent Q-codes and contains over 120 million records. Recognizing that missing Turkish descriptions in Wikidata would weaken the model's semantic representations, English descriptions were translated into Turkish using Meta AI's 3.3-billion parameter NLLB-200 machine translation model. This successfully increased the Turkish description coverage from 40.46\% to 94.57\%. Furthermore, the texts were linguistically enriched with morphological features, part-of-speech tags, and syntactic dependency tree analyses using Turkish.AI Natural Language Processing tools. Ultimately, TurkLink has become a massive resource containing roughly 590,000 articles, over 96 million words, more than 4.3 million entity mentions, and 15 million entity labels. Initial analyses using standard automatic detection tools showed that the overlap between Wikipedia hyperlinks and targeted entity boundaries remained at only 10.33\% due to abstract concepts and out-of-context links. To remove these noisy hyperlinks that would degrade the model's learning quality, a two-stage hybrid filtering mechanism was designed. In the first stage, rule-based filtering, a Breadth-First Search was conducted on Wikidata's ontological knowledge graph using the“instance of”(P31) and“subclass of”(P279) relations. This pre-filtered entities into 11 target classes (such as person, organization, facility, and artwork) based on carefully prepared whitelists and blacklists. Because the rule-based system could not fully analyze text context—leading to multi-class assignments and ambiguities—the second stage, model-based filtering, was introduced. In this stage, five modern Large Language Models (DeepSeek-R1 Distill, Gemma 3, Llama 3.1, Qwen3, and Trendyol LLM) were used to analyze the context in detail. The final, definitive entity classes were determined through a majority vote based on the outputs of these models. In the modeling phase, the ReFinED architecture was used as the foundation due to its high computational efficiency and zero-shot capabilities. Originally designed for English, this end-to-end architecture was integrated with the pre-trained“BERTurk 128k Cased”language model, which has a 128,000-word vocabulary, to better understand Turkish. The resulting new system was named“TR-ReFinED.”Given the agglutinative structure of Turkish, using a large-vocabulary and case-sensitive model that preserves word roots and suffix sequences is crucial. Additionally, to investigate the model's syntactic and morphological awareness, part-of-speech and dependency tags obtained from the TurkLink dataset were integrated directly into the model's embedding layer as 32-dimensional vectors. The performance of the developed models was comprehensively evaluated on both the TurkLink test set, which is reflecting in-domain features, and the cross-domain Mewsli-9 test Turkish subset, which is consisting of news texts to challenge the model's generalization ability. In Entity Disambiguation scenarios, where entity boundaries are given to the system in advance, the TR-ReFinED model achieved an F1-score of 96.46\% on the TurkLink set and 90.55\% on the Mewsli-9 set, demonstrating high success in resolving contextual ambiguities in Turkish texts. However, in end-to-end Entity Linking scenarios, where the model is also expected to predict entity boundaries, the system's performance dropped dramatically. Experiments revealed that, due to the structure of Turkish, end-to-end models face a serious bottleneck in determining entity boundaries. Derivational and inflectional suffixes attached directly to proper nouns cause millimeter-level shifts in boundary detection. Because strict evaluation algorithms consider these minor shifts entirely incorrect, the system's overall F1-score on the Mewsli-9 dataset fell to 0.2927. To overcome this structural weakness in the boundary detection phase, the Entity Linking problem was divided into two independent stages: the detection of entity mentions was delegated to a custom, separately trained BERTurk based Entity Mention Detection model, while TR-ReFinED was only responsible for linking these finalized candidates to the knowledge base. Separating the tasks and optimizing each model in its own area of expertise greatly improved overall system performance. The proposed pipeline architecture increased the F1-score to 0.5317 on the TurkLink test set and 0.4773 on the Mewsli-9 news texts. The success of this pipeline model became even more evident when compared to the literature's leading multilingual end-to-end Entity Linking models on the same Turkish Mewsli-9 test distribution. Faced with the complexity of Turkish, multilingual models (MD + mGENRE and mReFinED) produced a high volume of incorrect and unnecessary predictions, causing their precision to collapse. The best multilingual model, mReFinED, only reached an F1-score of 0.2960. In contrast, the pipeline model supported by external Entity Mention Detection, specifically designed for Turkish, achieved a strong balance of precision and recall. With an F1-score of 0.4773, it outperformed its strongest competitor by a wide margin of 0.1813 points. In conclusion, this thesis contributes TurkLink, which is a massive-scale resource enriched with morphological and semantic features, to the Turkish Natural Language Processing literature and presents a successful Entity Linking architecture for the Turkish language. The experiments demonstrate that for complex languages like Turkish, solving Entity Linking tasks using pipeline systems, where Entity Mention Detection and Entity Disambiguation processes feed into each other, is much more flexible, robust, and superior in performance compared to single end-to-end models.
Benzer Tezler
- Güney Dal'ın romanlarında varolma biçimleri ve göç olgusu
Existence forms and migration facts in Guney Dal's of novels
UTKU ÖZBAY
Yüksek Lisans
Türkçe
2017
Türk Dili ve EdebiyatıArdahan ÜniversitesiTürk Dili ve Edebiyatı Ana Bilim Dalı
DOÇ. DR. MİTAT DURMUŞ
- Semi-supervised learning based named entity recognition for morphologically rich languages
Morfolojik açıdan zengin dillerde yarı güdümlü öğrenme tekniğiyle varlık ismi tanıma
HAKAN DEMİR
Yüksek Lisans
İngilizce
2014
Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolBoğaziçi ÜniversitesiBilgisayar Mühendisliği Ana Bilim Dalı
YRD. DOÇ. ARZUCAN ÖZGÜR TÜRKMEN
- Aspect verbal en grec Istanbouliote: modele de description fonctionnaliste
Başlık çevirisi yok
EMİNE YAVAŞGEL
Yüksek Lisans
Fransızca
1993
Dilbilimİstanbul ÜniversitesiYabancı Diller Eğitimi Ana Bilim Dalı
PROF. DR. NÜKHET GÜZ
- Edebiyat sosyolojisi bağlamında toplumsal değerler ve değerler çatışması: Ahmed Arif, Faik Baysal ve Tarık Buğra örneği
Social values and value confliction in the context of sociology of literature: Ahmed Arif, Faik Baysal and Tarik Buğra sample
FADİME GÖKCEN YAŞAR
Yüksek Lisans
Türkçe
2023
SosyolojiKaramanoğlu Mehmetbey ÜniversitesiSosyoloji Ana Bilim Dalı
DR. ÖĞR. ÜYESİ NİLÜFER ÖZTÜRK AYKAÇ
- Ebû Şekur es-Salimî'nin et-Temhîd fi Beyâni Tevhid adlı eseri bağlamında İmam Eş'arî'ye ve Bidat mezheplere yönelik eleştirileri
Abu Shakur es-Salimî's criticisms about Imam Ash'ari and Bidat sects in the context of his work titled et-Tamhîd fi Beyâni Tawhid
İSMET KORKMAZ
Yüksek Lisans
Türkçe
2023
DinNecmettin Erbakan ÜniversitesiTemel İslam Bilimleri Ana Bilim Dalı
PROF. DR. RAMAZAN ALTINTAŞ