Ensemble QSAR Model for GSK3β Inhibitor Activity Prediction: Target-Specific Descriptors and PLIP Validation Junseo, Lee Advised by Prof. Kim, Mihyun Department of Pharmacy The Graduate School Gachon University Glycogen synthase kinase-3 beta represe...
Ensemble QSAR Model for GSK3β Inhibitor Activity Prediction: Target-Specific Descriptors and PLIP Validation Junseo, Lee Advised by Prof. Kim, Mihyun Department of Pharmacy The Graduate School Gachon University Glycogen synthase kinase-3 beta represents a therapeutically relevant target implicated in diverse pathological conditions including neurodegenerative diseases, mood disorders, and metabolic syndromes. Despite considerable research efforts, developing selective inhibitors with favorable drug-like properties remains challenging due to structural conservation across the kinase family. This investigation establishes a comprehensive computational framework for predicting GSK3β inhibitory activity through integration of machine learning methodologies with structure-based validation approaches. A curated dataset of 2,860 compounds with experimentally determined inhibitory activities was assembled from the ChEMBL database, with pChEMBL values serving as the continuous activity endpoint. Molecular characterization employed five distinct descriptor categories encompassing extended-connectivity fingerprints, conventional physicochemical descriptors, three-dimensional conformational features, Mordred descriptor suite, and target-specific pharmacophoric descriptors, yielding 1,763 initial features. Notably, eight GSK3β- specific descriptors underwent rigorous validation through protein-ligand interaction fingerprinting analysis using PLIP, establishing statistically significant associations with experimental binding modes observed in docked complexes (p < 0.05). Systematic ablation studies revealed that Mordred descriptors and ECFP fingerprints constitute the primary information sources, while structure-validated features provide mechanistically grounded enhancements. The minimum redundancy maximum relevance algorithm identified an optimal 200-feature subset that balanced predictive relevance with computational efficiency while preserving pharmacophore-validated descriptors. Model development employed heterogeneous ensemble learning combining six diverse base algorithms with a deep neural network meta-learner, with hyperparameters optimized through Bayesian optimization using the Optuna framework. Ensemble weights were determined via constrained least squares optimization rather than simple averaging, enabling adaptive emphasis on superior performers while maintaining diversity benefits. The final optimized model achieved test set performance of R² = 0.6498 and RMSE = 0.7366 pChEMBL units for continuous activity prediction, with classification accuracy of 73.4% for discrete activity categories. SHAP analysis provided instance-level interpretation of feature contributions, revealing that hydrogen bonding capacity, aromatic heterocycle content, and specific substructural motifs represent key molecular determinants of inhibitory potency. Applicability domain assessment through distance-based metrics and ensemble variance quantification identified compound regions where predictions exhibit enhanced reliability. This work demonstrates how thoughtful integration of diverse molecular representations, structure-based validation, sophisticated feature selection algorithms, and explainable artificial intelligence frameworks yields predictive models that are simultaneously accurate, interpretable, and mechanistically justified. The modeling strategies and validation approaches developed herein provide a generalizable template for QSAR investigations targeting other therapeutically important kinases, while specific insights regarding GSK3β inhibition offer actionable guidance for medicinal chemistry optimization campaigns. Keywords: QSAR modeling, GSK3-beta inhibitors, machine learning, ensemble methods, PLIP validation, feature selection, SHAP analysis, structure-based drug design