Learning cooperation in hunter-prey problem via state abstraction
Av avcı probleminde durum soyutlama yoluyla işbirliği öğrenme
- Tez No: 238621
- Danışmanlar: PROF. DR. FARUK POLAT
- Tez Türü: Yüksek Lisans
- Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
- Anahtar Kelimeler: Belirtilmemiş.
- Yıl: 2009
- Dil: İngilizce
- Üniversite: Orta Doğu Teknik Üniversitesi
- Enstitü: Fen Bilimleri Enstitüsü
- Ana Bilim Dalı: Bilgisayar Mühendisliği Bölümü
- Bilim Dalı: Bilgisayar Mühendisliği Ana Bilim Dalı
- Sayfa Sayısı: Belirtilmemiş.
Özet
Avcı av problemi Pekiştirmeli Öğrenme yöntemi için sıkça kullanılan bir deney alanıdır, ancak durum uzayı hacminin büyüklüğü ajan sayısına ve ortamın büyüklüğüne üstel bağlı olarak değişmektedir. Durum uzayının bu büyüklüğü standart Q-öğrenme algoritmasının kullanımını imkansız kıldığından, bu tez daha önce öğrenilmiş bilgiyi kullanıp daha büyük deney ortamlarında çalışabilen ajanlar üreterek, durum uzayı büyüklüğünün sabit tutmayı sağlayan bir yöntem tanıtmaktadır. Bu metot, Hiyerarşik Pekiştirmeli Öğrenme yöntemlerinden esinlenerek görevi daha basit alt görev seçimlerine bölen paralel alt görev mekanizmasından, bu yönteme yönelik bir durum gösterim tekniğinden ve bunun daha büyük ortamlar için genişletiminden oluşmaktadır. Deneysel sonuçlar önerilen yöntemin ortam parametrelerinden bağımsız, sabit büyüklükte bir durum uzayı kullanarak, el ile yazılmış algoritma kullanan ajanlara yakın, başarılı sonuçlar elde ettiğini göstermektedir.
Özet (Çeviri)
Hunter-Prey or Prey-Pursuit problem is a common toy domain for Reinforcement Learning, but the size of the state space is exponential in the parameters such as size of the grid or number of agents. As the size of the state space makes the flat Q-learning impossible to use for different scenarios, this thesis presents an approach to make the size of the state space constant by producing agents that use previously learned knowledge to perform on bigger scenarios containing more agents. Inspired from HRL methods, the method is composed of a parallel subtasks schema dividing the task into choices of simpler subtasks, a state representation technique convenient for this schema and its extension for bigger grids. Experimental results show that proposed method successfully provides agents that perform near to hand-coded agents by using constant sized state space independent from parameters of the domain.
Benzer Tezler
- Ameliyat öncesi hastaların ameliyata ilişkin duyguları, düşünceleri ve bilgi istekleri
The Pre-operative patiensts feelings throughts and information requirements concerning their surgical operations
KADRİYE BULDUKOĞLU
Yüksek Lisans
Türkçe
1987
HemşirelikCumhuriyet ÜniversitesiHemşirelik Ana Bilim Dalı
YRD. DOÇ. DR. MELİHA ATALAY
- Sivas ili cumhuriyet üniversitesi tip fakültesi numune ve sosyal sigortalari kurumu hastanelerinde çalişan hemşirelerin intra venöz (damar içi) sivi tedavisine ilişkin bilgi ve uygulamalari
The knowledge and practice related to the intravenous fluid theraph of the nurses working at the hospital of the medical faculty in cumhuriyet university, numune hastanesi and sosyal sigortalar kurumu hastanesi in sivas
İLKAY KALELİ
Yüksek Lisans
Türkçe
1985
HemşirelikCumhuriyet ÜniversitesiHemşirelik Ana Bilim Dalı
YRD. DOÇ. DR. MELİHA ATALAY