Geri Dön

Investigation of using the LSA model with similarity metrics for semantic-based web document clustering

Semantik bazlı web dokümanı kümelenmesi için benzeri metrikli LSA modelinin kullanımının incelenmesi

  1. Tez No: 492756
  2. Yazar: MASHHOOD ALI ALI
  3. Danışmanlar: YRD. DOÇ. DR. AYTUĞ BOYACI
  4. Tez Türü: Yüksek Lisans
  5. Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
  6. Anahtar Kelimeler: Metin madenciliği, Web tabanlı uygulamalar, Text mining, Web based applications
  7. Yıl: 2018
  8. Dil: İngilizce
  9. Üniversite: Fırat Üniversitesi
  10. Enstitü: Fen Bilimleri Enstitüsü
  11. Ana Bilim Dalı: Yazılım Mühendisliği Ana Bilim Dalı
  12. Bilim Dalı: Belirtilmemiş.
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Web belge kümelemesi, benzer web belgelerini, aynı kümedeki belgelerin diğer kümelerdeki belgelere göre semantik olarak daha yakın kategorize edildiği gruplar halinde bir araya getirmek için veri kümeleme tekniklerini kullanmaktadır. Belgeleri kümeleme yöntemlerinden biri, bu belgelerin içerdikleri konulara göre gruplandırılmasına dayanmaktadır. Konu tabanlı web belge kümeleme yönteminde kullanılan temel teknik, veri setinde bulunan terimler ve belgeler gibi her öğe için veri seti düzeyinde bir semantik (ör. konular) türeten ve LSA (Latent Semantic Analysis) olarak bilinen semantik analiz modelidir. LSA modeli literatürde, farklı şekillerde, varyasyonlarda ve farklı amaçlarda kullanılmıştır. Mevcut durumda LSA modelinin birçok kullanımı bulunduğundan, bu çalışmada, metin dokümanlarını semantik olarak kümelemede LSA modelinin en iyi şekilde kullanımı incelenmiştir. Bu sebeple, web belgelerinin kümelenmesinde en iyi performansı gösteren varyasyonu bulmak amacıyla LSA modelinin altı farklı semantik-benzerlik ölçümü ile kombinasyonları incelenmiştir. Metin kümelemesinde LSA modelini kullanımının en iyi varyasyonu, yine bu varyasyonun en çok kullanılan iki web dokümanı veri setine uygulanmasından sonra bulunmuştur. Sonuçlar aynı zamanda, web belge kümelemesi için LSA modelinin kullanımındaki her varyasyonun performansını göstermektedir.

Özet (Çeviri)

Web document clustering uses data clustering techniques to group similar web documents into groups, where the documents from the same cluster are more semantically similar than the documents in the other clusters. One of the methods of clustering the documents is based on the topics they contain. The main technique used for topic-based web document clustering is the using of a semantic-analysis model called Latent Semantic Analysis (LSA), which derives a corpus-level semantics (i.e. topics) for every element in the corpus such as, terms and documents. The LSA model has been used in the literature in different ways, variations and for different applications. In this study, we experimentally investigate the best use of the LSA model in semantically clustering the text documents, as there is more than one possible variation when one uses and implements the LSA model. To do so, we examined the LSA model in different combinations with six different semantic-similarity measures to find the best possible variation, which performs best in clustering web documents. The best variation of using the LSA model in text clustering was found after applying it to two commonly used web document datasets. The results also demonstrate the performance of each variation of using LSA model for the task of web document clustering.

Benzer Tezler

  1. Seri bağlı geri devirli aktif çamur sistemlerinde biyolojik kalıcı ürün oluşumunun modellenmesi

    Modelling of biological microbial product formation in activated sludge systems in series with recycle

    FEHİMAN ÇİNER

    Doktora

    Türkçe

    Türkçe

    1999

    Çevre Mühendisliğiİstanbul Teknik Üniversitesi

    PROF.DR. HASAN ALİ SAN

  2. Sabit yataklı kolon prosesi ile plastik atıklardan üretilmiş adsorbanları kullanarak sulu ortamlardan farmasötiklerin giderilmesinin incelenmesi

    Investigation of removal of pharmaceuticals from aqueous media using adsorbents produced from plastic waste by fixed bed column process

    GÜLSÜM ÖZÇELİK

    Doktora

    Türkçe

    Türkçe

    2026

    Kimya Mühendisliğiİstanbul Üniversitesi-Cerrahpaşa

    Kimya Mühendisliği Ana Bilim Dalı

    PROF. DR. SELİN ŞAHİN SEVGİLİ

    DOÇ. DR. EBRU KURTULBAŞ ŞAHİN

  3. Çeşitli doğal substratların yerel bir Aureobasidium pullulans suşunun pullulan üretimine etkilerinin incelenmesi

    Investigation of the effects of various natural substrates on the pullulan production by a domestic Aureobasidium pullulans strain

    BÜŞRA AKDENİZ

    Yüksek Lisans

    Türkçe

    Türkçe

    2019

    Gıda MühendisliğiHacettepe Üniversitesi

    Gıda Mühendisliği Ana Bilim Dalı

    PROF. DR. ZEKİYE YEŞİM ÖZBAŞ

  4. TFRS 9 standardı kapsamında karşılık uygulamalarının Türk bankacılık sektörüne etkisinin incelenmesi

    Investigation of the impact of provision applications on the Turkish banking sector within the scope of TFRS 9 standards

    TEVFİK AYDIN

    Yüksek Lisans

    Türkçe

    Türkçe

    2020

    BankacılıkGalatasaray Üniversitesi

    İşletme Ana Bilim Dalı

    DOÇ. DR. BANU DİNCER

  5. İki boyutlu sistemlerin magnetik özellikleri

    The Investigation of magnetic properties in two dimensional systems

    BAHADIR BOYACIOĞLU

    Doktora

    Türkçe

    Türkçe

    2001

    Fizik ve Fizik MühendisliğiAnkara Üniversitesi

    Fizik Ana Bilim Dalı

    DOÇ. DR. MESUDE SAĞLAM