Accurately predicting clinically relevant drug–drug interactions (DDIs) remains a persistent challenge due to the complexity of pharmacokinetic or pharmacodynamic mechanisms, sheer number of drug combinations to consider, and limited throughput of e...
Accurately predicting clinically relevant drug–drug interactions (DDIs) remains a persistent challenge due to the complexity of pharmacokinetic or pharmacodynamic mechanisms, sheer number of drug combinations to consider, and limited throughput of experiment-based DDI investigations. Although machine learning (ML)-driven approaches have been developed to address such limitations, these approaches have limited capacity for translating preclinical evidence and processing the high dimensionality of multimodal drug information. Moreover, prediction models trained on samples of drug combinations whose evidence of DDI is not known have limited applicability due to the issue of positive unlabeled data. Here, I present DDI-Expert, a Mixture-of-Experts (MoE)–based large language model framework designed to predict and explain the mechanisms of pharmacokinetic DDIs (PK DDIs), using heterogeneous multimodal drug features. To this end, using transfer learning on foundation models (i.e., BioClinical ModernBERT and T5), in which I inserted MoE architecture, I adapted the models to the prediction and explanation of PK DDI mechanisms, given multimodal drug characteristics of a drug pair. I further applied continual learning via Low-Rank Adaptation (LoRA) to incorporate human PK DDI evidence while mitigating catastrophic forgetting. I additionally employed conformal prediction to identify statistically reliable and minimally sufficient protein inputs, addressing the substantial noise introduced by drugs with a large number of associated features. DDI-Expert outperformed baseline models, including classical state-of-the-art machine learning approaches using structural information, LLMs without domain adaptation, and early-fusion multimodal models, across classification and explanation generation tasks. The MoE architecture enabled efficient multimodal integration with minimal overfitting, while the Noisy Top-K router ensured balanced expert utilization. Conformal prediction yielded compact, mechanistically relevant input subsets (approximately 3 proteins per drug) that improved in-context learning accuracy and reduced computational cost by up to 40-fold compared with using all inputs. LoRA-based continual learning surpassed retrieval-augmented generation (RAG) by providing robust domain-, class-, and task-incremental adaptation scenarios, enabling DDI-Expert to learn new human-specific labels and acquire an entirely new task—predicting AUC fold-change—without degrading prior knowledge. Model inspection revealed high explainability, strong semantic fidelity of generated explanations, and correct mechanistic reasoning even for less common PK processes (i.e., changes in the rate of excretion). These findings demonstrate that integrating MoE architecture, conformal prediction, and LoRA-based continual learning provides a scalable, interpretable, and clinically transferable framework for the prediction and explanation of PK DDIs in humans. These findings may enable a practical clinical decision-support tool for identifying high-risk drug combinations in patients on complex drug regimens. Moreover, this framework can enable early identification of DDI risks in humans and prioritization of DDI experimental studies in drug development.