Geri Dön

On the reinforcement learning analysis and learning the control of humanoid robot leg

Başlık çevirisi mevcut değil.

  1. Tez No: 400914
  2. Yazar: ÖNDER TUTSOY
  3. Danışmanlar: DR. MARTIN BROWN
  4. Tez Türü: Doktora
  5. Konular: Elektrik ve Elektronik Mühendisliği, Electrical and Electronics Engineering
  6. Anahtar Kelimeler: Belirtilmemiş.
  7. Yıl: 2013
  8. Dil: İngilizce
  9. Üniversite: The Unıversıty Of Manchester
  10. Enstitü: Yurtdışı Enstitü
  11. Ana Bilim Dalı: Belirtilmemiş.
  12. Bilim Dalı: Belirtilmemiş.
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Özet yok.

Özet (Çeviri)

Reinforcement learning is a method for learning sequential control actions or decisions using an instantaneous reward signal which implicitly defines a long term value function. It has been proposed to solve complex learning control problems without requiring explicit knowledge of the system's dynamics. Moreover, it has also been used as a model of cognitive learning in humans and applied to systems, such as humanoid robots, to study embodied cognition. However, there are relatively few results which describe the actual performance of such learning algorithms, even on relatively simple problems. In this thesis, simple test problems are used to investigate issues associated with the value function's representation and parametric convergence. In particular, the terminal convergence problem is analyzed with a known optimal (bang-bang) control policy where aim is to accurately learn the value function. For certain initial conditions, the closed form solution for the value function is calculated and it is shown to have a polynomial form. It is parameterized by terms which are functions of the unknown plant's parameters and the value function's discount factor and their convergence properties are analyzed. It is shown that the temporal difference error introduces a null space associated with the finite horizon basis function during the experiment. This is only non-singular when the experiment is terminated correctly and a number of (equivalent) solutions are described. It is also demonstrated that, in general, the test problem's dynamics are chaotic for random initial states and this causes a digital offset in the value function. Methods for estimating the offset are described and a dead-zone is proposed to switch off learning in the chaotic region. Another value function estimation test problem is then proposed which uses a saturated piecewise linear control signal. This is a more realistic control scenario and it is also shown to address the chaotic dynamics problem. It is shown that the condition of the learning problem depends on both the saturation threshold and the value function's discount factor and that a badly conditioned learning problem may result. Moreover, it is proved that the temporal difference error introduce a trajectory null space associated with the differenced higher order bases until the saturation threshold of the saturated piecewise linear control signal. These results are then used to explain the behaviour of reinforcement learning algorithms when higher order systems are used and the impact of function approximation algorithms and exploration noise is discussed. Finally, a central pattern generator based reinforcement learning algorithm is applied to a single leg of a robot where the target is to generate appropriate control signals for each joint.

Benzer Tezler

  1. Explorations on inverse reinforcement learning for the analysis of motor control and cognitive decision making mechanisms of the brain

    Motor kontrol ve beynin bilişsel karar verme mekanizmalarını analiz etmek üzere tersine pekiştirmeli öğrenme ile keşifler

    EMİR ARDİTİ

    Yüksek Lisans

    İngilizce

    İngilizce

    2021

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolÖzyeğin Üniversitesi

    Bilgisayar Bilimleri Ana Bilim Dalı

    PROF. DR. ERHAN ÖZTOP

    DR. ÖĞR. ÜYESİ REYHAN AYDOĞAN

    DOÇ. DR. EMRE UĞUR

  2. Nesnelerin internetinde derin öğrenmeye dayalı veri analizi ve bilgi çıkarımı

    Deep learning based data analysis and information extraction in the internet of things

    İBRAHİM KÖK

    Doktora

    Türkçe

    Türkçe

    2020

    Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve KontrolGazi Üniversitesi

    Bilgisayar Bilimleri Ana Bilim Dalı

    PROF. DR. SUAT ÖZDEMİR

  3. Ses eğitimi dersinde harmanlanmış öğrenmenin akademik başarıya ve ses performansına etkisi

    The effect of blended learning on academic achievement and vocal performance in the vocal training course

    BATUHAN TÜFEKÇİ

    Doktora

    Türkçe

    Türkçe

    2026

    Eğitim ve ÖğretimGazi Üniversitesi

    Güzel Sanatlar Eğitimi Ana Bilim Dalı

    DR. ÖĞR. ÜYESİ MURAT KARABULUT

  4. İlkokul matematik dersinde ters yüz öğrenme destekli oyunlaştırılmış akran öğretiminin ders başarısı, motivasyonu ve öğrencilerin sosyal becerileri üzerindeki etkisi

    The effect of flip learning supported gamified peer teaching on students' achievement, motivation and social skills in primary school mathematics course

    HARUN ASLAN

    Yüksek Lisans

    Türkçe

    Türkçe

    2025

    Eğitim ve ÖğretimBartın Üniversitesi

    Eğitim Bilimleri Ana Bilim Dalı

    DOÇ. DR. MUSTAFA FİDAN

  5. Sosyal bilgiler öğretiminde öğrenme stiline göre yapılan eğitim faaliyetlerinin öğrencilerin kavramsal gelişim sürecine etkisi

    Impact of educational activities applied in accordance with learning styles on the conceptual development process of students in teaching social studies

    ÖZGE PAŞAOĞLU KOÇAK

    Yüksek Lisans

    Türkçe

    Türkçe

    2019

    Eğitim ve ÖğretimYıldız Teknik Üniversitesi

    Temel Eğitim Ana Bilim Dalı

    DOÇ. DR. MUSTAFA ŞEKER