Abstract
Immune cell migration from lymphoid organs to inflamed tissues is regulated by cell-surface receptors sensitive to environmental changes. G protein-coupled receptors involved in immune cell trafficking have been used in the design of anti-inflammatory drugs providing large datasets of active compounds deposited in e.g. ChEMBL. Here, we evaluated these datasets in terms of their applicability to machine learning (ML) aimed at compound activity prediction. Mixed-receptor datasets selected based on sequence homology were used in training of deep neural networks (DNNs) and gradient boosting machines (GBMs) to develop accurate ML classifiers for receptors with sparse ligand datasets in ChEMBL such as CCR6, CCR7, CCR8, CCR9, CCR10, CXCR5, and CCRL2. This ligand-based approach to drug design was compared with structure-based virtual screening (SBVS) using refined, ready-for-VS, structural models of immune cell trafficking receptors recently implemented in GPCRVS. The strengths and weaknesses of DNNs, GBMs, and SBVS were described in terms of applicability of various data schemes. As a result, a comprehensive, multi-level approach based on publicly accessible repositories has been proposed to facilitate immunomodulatory drug design.