Transcriptome dynamically captures cellular states in response to intrinsic and extrinsic factors and offers critical insights into disease biology and therapeutic potential. The expanding availability of large-scale transcriptomic datasets, such as T...
Transcriptome dynamically captures cellular states in response to intrinsic and extrinsic factors and offers critical insights into disease biology and therapeutic potential. The expanding availability of large-scale transcriptomic datasets, such as The Cancer Genome Atlas (TCGA) and Library of Integrated Network-based Cellular Signatures (LINCS) L1000, has increased interest in leveraging transcriptomic information to enhance clinical outcome predictions and facilitate phenotype-driven drug discovery.
However, fully utilizing transcriptomic data remains challenging due to its inherent context dependency and the intricate gene-gene interactions underlying biological phenotypes. Moreover, integrating transcriptomic profiles with heterogeneous clinical data or using them to guide molecular generation requires computational models that are both expressive and biologically interpretable. Transformer-based architectures are characterized by context-sensitive representation learning and attention-driven interpretability and offer promising avenues for addressing these challenges. However, their specific applications to biomedical domains remain relatively unexplored.
This thesis tackles two central problems in transcriptome-driven biomedical modeling. First, Transcriptome Transformer (TxT), a multi-task Transformer model, is introduced that utilizes transcriptomic data to predict patient survival outcomes, while clinical features are predicted as auxiliary tasks to improve learning of transcriptomic patterns. TxT employs multi-head attention mechanisms to explicitly model gene-gene interactions, complemented by a novel transcriptome-based positional embedding strategy, which significantly improves predictive accuracy while maintaining interpretability. Validation across multiple cancer datasets illustrates TxT’s capability to elucidate biologically meaningful gene contributions to clinical prognoses.
Second, a generative model, GGIFragGPT, was developed as a molecular generation framework conditioned on biologically informed gene-level embeddings derived from Geneformer, a Transformer model pre-trained on approximately 95 million single-cell transcriptomes. Gene-level embeddings generated by Geneformer from transcriptomic perturbation signatures serve as conditions for a fragment-based autoregressive molecular generation process, enabling GGIFragGPT to reliably produce chemically valid and phenotypically relevant compounds. Case studies, such as the generation of CDK7-specific molecules guided by shRNA-induced expression signatures, demonstrate the model's capability to yield molecules closely aligned with targeted biological phenotypes, with interpretability facilitated by attention-based analysis.
Collectively, these contributions establish a comprehensive computational framework for leveraging transcriptomics in biomedical prediction and phenotypic molecule design. By combining Transformer-based approaches with explicit representations of gene–gene interactions, clinical variables, and transcriptomic perturbations, this thesis presents interpretable and biologically grounded frameworks that support progress in data-driven medicine and therapeutic discovery.