Machine learning applications in drug discovery and development
İlaç keşfi ve tedavisinde yapay zeka uygulamaları
- Tez No: 985779
- Danışmanlar: DR. SHAN HE
- Tez Türü: Doktora
- Konular: Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control
- Anahtar Kelimeler: Akılcı ilaç kullanımı, Farmasötik kimya, Protein bağlama, Yapay zeka ve makine öğrenmesi dersi, Rational drug use, Chemistry-pharmaceutical, Protein binding, Artificial intelligence and machine learning course
- Yıl: 2025
- Dil: İngilizce
- Üniversite: The Unıversıty Of Bırmıngham
- Enstitü: Yurtdışı Enstitü
- Ana Bilim Dalı: Mühendislik ve Doğa Bilimleri Ana Bilim Dalı
- Bilim Dalı: Bilgisayar Bilimleri Bilim Dalı
- Sayfa Sayısı: Belirtilmemiş.
Özet
İlaç keşfi ve geliştirilmesinde makine öğrenmesi (ML) uygulamaları, geleneksel yöntemlere kıyasla daha hızlı sonuçlar, artırılmış doğruluk ve maliyetlerin azaltılması gibi önemli avantajlar sunmaktadır. Ancak, daha yüksek başarı elde etme çabası, derin öğrenme (DL) yaklaşımlarını içeren karmaşık modellerin benimsenmesine neden olmuş ve bu da yorumlanabilirliği olumsuz etkilemiştir. Bu durum, performans ile yorumlanabilirliği dengeleyen alternatif yaklaşımlara duyulan ihtiyacı artırmıştır. Basit ama etkili yöntemler, ilaç keşfi ve geliştirme süreçlerinin daha iyi anlaşılmasını sağlamak için gereklidir. Özellikle, aşağıdaki üç kritik alan bu bağlamda ele alınmıştır: 1. Blind docking (körü körüne kenetlenme) 2. Allosterik bağlanma bölgelerinin tespiti 3. PROteolysis TArgeting Chimaeras (PROTAC) taraması (i) Blind Docking ve Bağlanma Bölgesi Tahmini Protein yüzeyinin analiz edilerek küçük bir molekülün bağlanma bölgesinin ve bağlanma afinitesinin tahmin edilmesi, ilaç keşfinde kritik ancak zorlu bir görevdir. Blind docking, bağlanma bölgelerini rastgele seçerek protein yüzeyinde kenetlenme gerçekleştirmektedir. Ancak yerel (lokal) docking'e kıyasla daha az doğruluklu ve güvenilir bir yöntemdir, çünkü geniş arama alanı yeterince iyi örneklenememektedir. Bu sorunu gidermek için, bazı çalışmalar bağlanma boşluğu (cavity) tespiti yapan araçları kullanarak blind docking'in doğruluğunu artırmıştır. Ancak, bu yöntemlerin başarısı büyük ölçüde kullanılan cavity tespit aracının kalitesine bağlıdır. Bu bağımlılığı azaltmak için, CoBDock adlı yeni bir blind, paralel docking yöntemi geliştirilmiştir. CoBDock, makine öğrenmesi algoritmalarını kullanarak hem docking hem de cavity tespiti sonuçlarını entegre eder ve böylece bağlanma bölgesi tespitini ve bağlanma pozisyonu doğruluğunu artırır. Yapılan deneylerde, PDBBind 2020, ADS, MTi, DUD-E veCASF-2016 veri kümelerinde CoBDock'un, diğer güncel cavity tespit araçları ve blind docking yöntemlerinden daha başarılı olduğu gösterilmiştir. (ii) Allosterik Bağlanma Bölgelerinin Tespiti Allosterik modülatörler, ortosterik ligandlara kıyasla artırılmış seçicilik ve doygunluk gibi avantajlar sunmaktadır. Yeni allosterik bölgelerin tespit edilmesi, yenilikçi ilaçların geliştirilmesi açısından önemli fırsatlar sunarken, biyolojik mekanizmaların daha iyi anlaşılmasını da sağlamaktadır. Mevcut ML tabanlı yöntemlerden PASSer, sadece 3D yapısal veriye dayalı olarak allosterik bölgeleri belirlemede sınırlı başarı göstermektedir. Bu nedenle, amino asit bazlı destekleyici bilgilerin 3D yapısal veriye entegre edilmesi daha yüksek doğruluk ve sağlamlık sağlayabilir. Bu amaçla, 9460 farklı ve çeşitli özellik içeren bir veri seti oluşturularak, yeni ve güçlü bir model olan MEF-AlloSite geliştirilmiştir. Bu model, küçük eğitim setleriyle (yalnızca 90 protein) çalışacak şekilde optimize edilmiştir ve en iyi özellikleri seçerek tahmin performansını artırmaktadır. MEF-AlloSite, PASSer2.0 ve PASSerRank gibi güncel yöntemlerle karşılaştırılmıştır. Üç farklı test kümesiyle yapılan 51 farklı deneyde, Student's t-testi ve Cohen's D analizleri kullanılarak ortalama hassasiyet ve ROC AUC skorları değerlendirilmiştir. MEF-AlloSite'un ortalama hassasiyet ve ROC AUC skorlarının, diğer yöntemlere kıyasla %1-6 daha yüksek olduğu ve bu farkın istatistiksel olarak anlamlı olduğu gösterilmiştir. (iii) PROTAC'lar İçin Modelleme ve Sıralama Optimizasyonu PROTAC'lar, hedef proteine E3 ligaz bağlanmasını sağlayarak proteolizi tetikleyen küçük moleküllerdir. Bu yöntem, yüksek etkinlik, düşük doz gereksinimi ve geleneksel yöntemlerle ilaçlanamayan hedefler üzerinde etkili olma gibi avantajlar sunmaktadır. Ancak, PROTAC keşfi, bağlayıcı yapıları optimize etmek için çok sayıda analoğun sentezlenmesini gerektirir. Doğru bir üçlü kompleks (ternary complex) oluşumunu tahmin etmek zordur ve deneme-yanılma süreçleri uzun sürebilir. Bu problemi çözmek için, MEGA PROTAC adlı yeni bir yöntem geliştirilmiştir. MEGA PROTAC, protein-protein kompleksleri (PPC) için MEGADOCK'u kullanarak docking işlemi gerçekleştirmekte ve bir ön araştırma alanı belirlemektedir. Daha sonra, sıralama birleştirme (rank aggregation) ve ızgara arama (grid search) stratejileri kullanılarak en uygun PPC'ler seçilmektedir. MEGA PROTAC, deneysel olarak doğrulanmış 22 yapıda test edilmiştir ve mevcut en iyi yöntem olan Bayesian Optimization for Ternary Complex Prediction (BOTCP) ile karşılaştırılmıştır.MEGA PROTAC, 22 test vakasının 16'sında BOTCP'den daha iyi performans göstermiştir ve DockQ skorunda %18 daha yüksek ortalama ve %35 daha yüksek medyan değer elde etmiştir. Ayrıca, MEGA PROTAC'ın, BOTCP'ye kıyasla %75 daha iyi sıralama sağladığı ve daha az kümeleme ile en iyi DockQ skorunu bulduğu belirlenmiştir. Son olarak, MEGA PROTAC, BOTCP'ye kıyasla iki kat daha hızlı bir şekilde kabul edilebilir DockQ skorlarına ulaşmıştır ve bulunan kümelerde daha fazla doğal yapıya yakın yapı içermektedir. Sonuç Bu tez, blind docking, allosterik bağlanma bölgelerinin tespiti ve PROTAC keşfi gibi ilaç keşfi açısından önemli üç temel problemi ele almaktadır. Çalışmalar, makine öğrenmesi yöntemlerinin ilaç keşfinde etkinliğini artırabileceğini ve mevcut yöntemlere kıyasla daha yüksek doğruluk sağlayabileceğini göstermektedir. Özellikle CoBDock, MEF-AlloSite ve MEGA PROTAC modelleri, performans ve yorumlanabilirliği optimize eden yenilikçi yaklaşımlar sunmaktadır. Elde edilen sonuçlar, bu modellerin endüstri ve akademik araştırmalarda önemli bir potansiyele sahip olduğunu göstermektedir.
Özet (Çeviri)
Machine learning (ML) applications in drug discovery and development offer sig- nificant advantages over traditional methods, such as faster outcomes, improved ac- curacy, and reduced costs. However, the drive to achieve enhanced performance has often resulted in adopting highly complex models, including deep learning (DL) ap- proaches, which can compromise interpretability. This has highlighted the need for alternative approaches that balance performance and interpretability. Specifically, straightforward yet effective methods are essential for advancing our understand- ing of key drug discovery and development processes. Such approaches should address performance limitations without sacrificing clarity, particularly in critical areas such as (i) blind docking, (ii) the identification of allosteric binding sites, and (iii) PROteolysis TArgeting Chimaeras (PROTAC) screening. (i) Probing the surface of proteins to predict the binding site and binding affinity for a given small molecule is a critical but challenging task in drug discovery. Blind docking addresses this issue by performing docking on binding regions randomly sampled from the entire protein surface. However, compared with local docking, blind docking is less accurate and reliable because the docking space is too ample to be sufficiently sampled. Cavity detection-guided blind docking methods improved the accuracy by using cavity detection (also known as binding site detection) tools to guide the docking procedure. However, it is worth noting that the performance of these methods heavily relies on the quality of the cavity detection tool. This con- straint, namely the dependence on a single cavity detection tool, significantly im- pacts the overall performance of cavity detection-guided methods. To overcome this limitation, we proposed Consensus Blind Dock (CoBDock), a novel blind, parallel docking method that uses ML algorithms to integrate docking and cavity detection results to improve not only binding site identification but also pose prediction ac- curacy. Our experiments on several datasets, including PDBBind 2020, ADS, MTi, DUD-E, and CASF-2016, showed that CobDock has better binding site and bind- ing mode performance than other state-of-the-art cavity detector tools and blind docking methods. (ii) A crucial mechanism for controlling the actions of proteins is allostery. Al- losteric modulators have the potential to provide many benefits in comparison to orthosteric ligands, such as increased selectivity and saturability of their effect. Iden- tifying new allosteric sites presents prospects for creating innovative medications and enhances our understanding of fundamental biological mechanisms. Allosteric sites are increasingly found in different protein families through various techniques, such as ML applications, which opens up possibilities for creating completely novel medications with diverse chemical structures. ML methods, such as PASSer, exhibit limited efficacy in accurately finding allosteric binding sites when relying solely on 3D structural information. Prior to conducting feature selection for allosteric bind- ing site identification, integration of supporting amino-acid-based information to 3D structural knowledge is advantageous. This approach can enhance performance by ensuring accuracy and robustness. Therefore, we have developed an accurate and robust model called Multimodel Ensemble Feature Selection for Allosteric Site Iden- tification (MEF-AlloSite) after collecting 9460 relevant and diverse features from the literature to characterize pockets. The model employs an accurate and robust multi- model feature selection technique for the small training set size of only 90 proteins to improve predictive performance. This state-of-the-art technique increased the performance in allosteric binding site identification by selecting promising features from 9460 features. Also, the relationship between selected features and allosteric binding sites enlightened the understanding of complex allostery for proteins by analyzing chosen features. MEF-AlloSite and state-of-the-art allosteric site identifi- cation methods such as PASSer2.0 and PASSerRank have been tested on three test cases 51 times with a different split of the training set. The Student's t-test and Co- hen's D value have been used to evaluate the average precision and ROC AUC score distribution. On three test cases, most of the p-values (< 0.05) and the majority of Cohen's D values (> 0.5) showed that MEF-AlloSite's 1-6% higher mean of average precision and ROC AUC than state-of-the-art allosteric site identification methods are statistically significant. (iii) Proteolysis-targeting chimeras (PROTACs), which induce proteolysis by re- cruiting an E3 ligase to dock into a target protein, are acquiring popularity as a novel pharmacological modality because of unique features of PROTAC, including high potency, low dosage, effectiveness on undruggable targets. While PROTACs are promising prospects as chemical probes and therapeutic agents, their discovery usually necessitates the synthesis of numerous analogs to explore variations on the chemical linker structure exhaustively. Without extensive trial and error, it is un- known how to link the two protein-recruiting moieties to facilitate the formation of a productive ternary complex. Although molecular docking-based and optimization pipelines have been designed to predict ternary complexes, guiding rational PRO- TAC design, they have suffered from limited predictive performance in the quality of the ternary structure and their ranks. Therefore, MEGA PROTAC has been designed to enhance the performance in the quality and ranking of ternary structures. MEGA PROTAC employs MEGADOCK to execute docking for protein-protein complexes (PPCs). The docking establishes an initial exploration area for PPCs. A sequential filtration strategy combined with rank aggregation is employed to choose a subset of PPCs for grid search. Once candidate PPCs are selected, a grid search method is used separately for translation and rotation. The remaining proteins have been grouped into clusters, and MEGA PROTAC further filters these clusters based on the energy score of the proteins within each cluster. MEGA PROTAC utilizes rank aggrega- tion to choose the best clusters and then employs MEGADOCK to dock PROTAC into the selected PPCs, forming a ternary structure. Finally, MEGA PROTAC was tested on 22 experimentally validated structures representing all currently available data. These cases were used to compare MEGA PROTAC with the state-of-the-art method, Bayesian Optimization for Ternary Complex Prediction (BOTCP). MEGA PROTAC outperformed BOTCP on 16 test cases out of 22 cases, achieving a higher maximum DockQ score with an 18% higher mean and 35% higher median. Also, MEGA PROTAC exhibited 75% superior ranks and a reduced cluster number for maximum DockQ score compared to BOTCP. Also, MEGA PROTAC outperforms BOTCP by achieving a twofold improvement in locating the first acceptable DockQ scores, with a more significant proportion of near-native structures within the de- tected cluster.
Benzer Tezler
- Makine öğrenme yaklaşımlarının biyoinformatikte ilaç geliştirme probleminde kullanılması
Using machine learning approaches in drug development problem in bioinformatics
TUĞÇE SEMERCİ
Yüksek Lisans
Türkçe
2023
İstatistikHacettepe Üniversitesiİstatistik Ana Bilim Dalı
PROF. DR. ÇAĞDAŞ HAKAN ALADAĞ
- Reinforcement learning-driven ensemble neural networks for heart disease prediction
Kalp hastalığı tahmini için takviyeli öğrenme tabanlı topluluk sinir ağları
ÖZGE HÜSNİYE NAMLI DAĞ
Doktora
İngilizce
2025
Endüstri ve Endüstri Mühendisliğiİstanbul Teknik ÜniversitesiEndüstri Mühendisliği Ana Bilim Dalı
DOÇ. DR. SEDA YANIK ÖZBAY
- Efficient optimization algorithms for computational biology
Hesaplamalı biyolojide etkin eniyileme algoritmaları
OĞUZ CAN BİNATLI
Doktora
İngilizce
2024
Endüstri ve Endüstri MühendisliğiKoç ÜniversitesiEndüstri Mühendisliği ve Operasyon Yönetimi
PROF. DR. MEHMET GÖNEN
- Çok değişkenli yapay sinir ağı operatörlerinin yaklaşım teorisindeki uygulamaları
Applications of multivariable artifical neural network operators in approximation theory
CANDAŞ DİNÇ
Yüksek Lisans
Türkçe
2025
MatematikAnkara ÜniversitesiMatematik Ana Bilim Dalı
PROF. DR. FATMA TAŞDELEN YEŞİLDAL
- Computational approaches to study drug resistance mechanisms
İlaç direnç mekanizmaları için işlemsel yaklaşımlar
ZOYA KHALID
Doktora
İngilizce
2017
BiyolojiSabancı ÜniversitesiMoleküler Biyoloji, Genetik ve Biyoteknoloji Ana Bilim Dalı
Prof. Dr. İSMAİL ÇAKMAK