RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Leveraging Transcriptomics with Transformer-based Deep Learning for Prognosis Prediction and Phenotype-driven Molecule Generation = 전사체 정보를 활용한 트랜스포머 기반 환자 예후 예측 및 표현형 기반 분자 생성

    한글로보기

    https://www.riss.kr/link?id=T17450560

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Transcriptome dynamically captures cellular states in response to intrinsic and extrinsic factors and offers critical insights into disease biology and therapeutic potential. The expanding availability of large-scale transcriptomic datasets, such as The Cancer Genome Atlas (TCGA) and Library of Integrated Network-based Cellular Signatures (LINCS) L1000, has increased interest in leveraging transcriptomic information to enhance clinical outcome predictions and facilitate phenotype-driven drug discovery.

    However, fully utilizing transcriptomic data remains challenging due to its inherent context dependency and the intricate gene-gene interactions underlying biological phenotypes. Moreover, integrating transcriptomic profiles with heterogeneous clinical data or using them to guide molecular generation requires computational models that are both expressive and biologically interpretable. Transformer-based architectures are characterized by context-sensitive representation learning and attention-driven interpretability and offer promising avenues for addressing these challenges. However, their specific applications to biomedical domains remain relatively unexplored.

    This thesis tackles two central problems in transcriptome-driven biomedical modeling. First, Transcriptome Transformer (TxT), a multi-task Transformer model, is introduced that utilizes transcriptomic data to predict patient survival outcomes, while clinical features are predicted as auxiliary tasks to improve learning of transcriptomic patterns. TxT employs multi-head attention mechanisms to explicitly model gene-gene interactions, complemented by a novel transcriptome-based positional embedding strategy, which significantly improves predictive accuracy while maintaining interpretability. Validation across multiple cancer datasets illustrates TxT’s capability to elucidate biologically meaningful gene contributions to clinical prognoses.

    Second, a generative model, GGIFragGPT, was developed as a molecular generation framework conditioned on biologically informed gene-level embeddings derived from Geneformer, a Transformer model pre-trained on approximately 95 million single-cell transcriptomes. Gene-level embeddings generated by Geneformer from transcriptomic perturbation signatures serve as conditions for a fragment-based autoregressive molecular generation process, enabling GGIFragGPT to reliably produce chemically valid and phenotypically relevant compounds. Case studies, such as the generation of CDK7-specific molecules guided by shRNA-induced expression signatures, demonstrate the model's capability to yield molecules closely aligned with targeted biological phenotypes, with interpretability facilitated by attention-based analysis.

    Collectively, these contributions establish a comprehensive computational framework for leveraging transcriptomics in biomedical prediction and phenotypic molecule design. By combining Transformer-based approaches with explicit representations of gene–gene interactions, clinical variables, and transcriptomic perturbations, this thesis presents interpretable and biologically grounded frameworks that support progress in data-driven medicine and therapeutic discovery.
    번역하기

    Transcriptome dynamically captures cellular states in response to intrinsic and extrinsic factors and offers critical insights into disease biology and therapeutic potential. The expanding availability of large-scale transcriptomic datasets, such as T...

    Transcriptome dynamically captures cellular states in response to intrinsic and extrinsic factors and offers critical insights into disease biology and therapeutic potential. The expanding availability of large-scale transcriptomic datasets, such as The Cancer Genome Atlas (TCGA) and Library of Integrated Network-based Cellular Signatures (LINCS) L1000, has increased interest in leveraging transcriptomic information to enhance clinical outcome predictions and facilitate phenotype-driven drug discovery.

    However, fully utilizing transcriptomic data remains challenging due to its inherent context dependency and the intricate gene-gene interactions underlying biological phenotypes. Moreover, integrating transcriptomic profiles with heterogeneous clinical data or using them to guide molecular generation requires computational models that are both expressive and biologically interpretable. Transformer-based architectures are characterized by context-sensitive representation learning and attention-driven interpretability and offer promising avenues for addressing these challenges. However, their specific applications to biomedical domains remain relatively unexplored.

    This thesis tackles two central problems in transcriptome-driven biomedical modeling. First, Transcriptome Transformer (TxT), a multi-task Transformer model, is introduced that utilizes transcriptomic data to predict patient survival outcomes, while clinical features are predicted as auxiliary tasks to improve learning of transcriptomic patterns. TxT employs multi-head attention mechanisms to explicitly model gene-gene interactions, complemented by a novel transcriptome-based positional embedding strategy, which significantly improves predictive accuracy while maintaining interpretability. Validation across multiple cancer datasets illustrates TxT’s capability to elucidate biologically meaningful gene contributions to clinical prognoses.

    Second, a generative model, GGIFragGPT, was developed as a molecular generation framework conditioned on biologically informed gene-level embeddings derived from Geneformer, a Transformer model pre-trained on approximately 95 million single-cell transcriptomes. Gene-level embeddings generated by Geneformer from transcriptomic perturbation signatures serve as conditions for a fragment-based autoregressive molecular generation process, enabling GGIFragGPT to reliably produce chemically valid and phenotypically relevant compounds. Case studies, such as the generation of CDK7-specific molecules guided by shRNA-induced expression signatures, demonstrate the model's capability to yield molecules closely aligned with targeted biological phenotypes, with interpretability facilitated by attention-based analysis.

    Collectively, these contributions establish a comprehensive computational framework for leveraging transcriptomics in biomedical prediction and phenotypic molecule design. By combining Transformer-based approaches with explicit representations of gene–gene interactions, clinical variables, and transcriptomic perturbations, this thesis presents interpretable and biologically grounded frameworks that support progress in data-driven medicine and therapeutic discovery.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    전사체(transcriptome)는 내·외부 자극에 따른 세포 상태를 동적으로 반영하며, 질병의 생물학적 기전 이해와 치료 전략 탐색에 중요한 단서를 제공한다. The Cancer Genome Atlas (TCGA) 및 Library of Integrated Network-based Cellular Signatures (LINCS) L1000과 같은 대규모 전사체 데이터의 축적은 전사체 정보를 활용한 임상 예후 예측과 표현형 기반 약물 발굴 연구에 대한 관심을 크게 증대시켰다.

    그러나 전사체 데이터는 생물학적 맥락 의존성이 강하고, 표현형을 결정하는 유전자 간 상호작용이 복잡하여 이를 효과적으로 활용하는 데 한계가 존재한다. 또한 전사체 프로파일을 이질적인 임상 데이터와 통합하거나 분자 생성 과정의 조건으로 활용하기 위해서는, 높은 표현력을 가지면서도 생물학적으로 해석 가능한 계산 모델이 요구된다. Transformer 기반 모델은 문맥 인식 표현 학습과 attention 메커니즘을 통한 해석 가능성을 특징으로 하여 이러한 문제를 해결할 수 있는 잠재력을 지니고 있으나, 바이오의학 분야에서의 구체적인 적용 사례는 아직 제한적인 상황이다.

    본 학위 논문에서는 전사체 기반 바이오의학 모델링에서의 두 가지 핵심 문제를 다룬다. 첫째, 전사체 데이터를 활용하여 환자의 생존 예후를 예측하는 다중 과제 Transformer 모델인 Transcriptome Transformer (TxT)를 제안한다. TxT는 생존 예측을 주 과제로 설정하고, 임상 변수를 보조 과제로 함께 학습함으로써 전사체 패턴에 대한 표현 학습을 향상시킨다. 다중 헤드 attention 메커니즘을 통해 유전자–유전자 상호작용을 명시적으로 모델링하며, 전사체 특성을 반영한 새로운 positional embedding 기법을 도입하여 예측 성능을 향상시키는 동시에 모델의 해석 가능성을 유지한다. 다수의 암 코호트에 대한 검증을 통해 TxT가 임상 예후와 연관된 생물학적으로 의미 있는 유전자 기여를 효과적으로 포착할 수 있음을 확인하였다.

    둘째, 단일세포 전사체 약 9천5백만 건으로 사전 학습된 Transformer 모델인 Geneformer로부터 도출한 유전자 수준 임베딩을 조건으로 사용하는 분자 생성 모델 GGIFragGPT를 제안한다. 전사체 교란 시그니처로부터 생성된 유전자 임베딩을 기반으로 fragment 단위의 자기회귀적 분자 생성 과정을 수행함으로써, 화학적으로 타당하고 표현형과 연관된 분자를 안정적으로 생성할 수 있다. shRNA 유도 발현 시그니처를 활용한 CDK7 표적 분자 생성 사례를 통해, 제안한 모델이 목표 생물학적 표현형과 정합성이 높은 분자를 생성할 수 있음을 보였으며, attention 기반 분석을 통해 생성 과정의 해석 가능성 또한 제시하였다.

    종합적으로, 본 연구는 전사체 데이터를 임상 예측과 표현형 기반 분자 설계에 효과적으로 활용하기 위한 통합적 계산 프레임워크를 제시한다. Transformer 기반 모델에 유전자–유전자 상호작용, 임상 변수, 전사체 교란 정보를 명시적으로 통합함으로써, 데이터 기반 정밀의료 및 치료제 개발을 위한 해석 가능하고 생물학적으로 타당한 접근법을 제공한다.
    번역하기

    전사체(transcriptome)는 내·외부 자극에 따른 세포 상태를 동적으로 반영하며, 질병의 생물학적 기전 이해와 치료 전략 탐색에 중요한 단서를 제공한다. The Cancer Genome Atlas (TCGA) 및 Library of Integr...

    전사체(transcriptome)는 내·외부 자극에 따른 세포 상태를 동적으로 반영하며, 질병의 생물학적 기전 이해와 치료 전략 탐색에 중요한 단서를 제공한다. The Cancer Genome Atlas (TCGA) 및 Library of Integrated Network-based Cellular Signatures (LINCS) L1000과 같은 대규모 전사체 데이터의 축적은 전사체 정보를 활용한 임상 예후 예측과 표현형 기반 약물 발굴 연구에 대한 관심을 크게 증대시켰다.

    그러나 전사체 데이터는 생물학적 맥락 의존성이 강하고, 표현형을 결정하는 유전자 간 상호작용이 복잡하여 이를 효과적으로 활용하는 데 한계가 존재한다. 또한 전사체 프로파일을 이질적인 임상 데이터와 통합하거나 분자 생성 과정의 조건으로 활용하기 위해서는, 높은 표현력을 가지면서도 생물학적으로 해석 가능한 계산 모델이 요구된다. Transformer 기반 모델은 문맥 인식 표현 학습과 attention 메커니즘을 통한 해석 가능성을 특징으로 하여 이러한 문제를 해결할 수 있는 잠재력을 지니고 있으나, 바이오의학 분야에서의 구체적인 적용 사례는 아직 제한적인 상황이다.

    본 학위 논문에서는 전사체 기반 바이오의학 모델링에서의 두 가지 핵심 문제를 다룬다. 첫째, 전사체 데이터를 활용하여 환자의 생존 예후를 예측하는 다중 과제 Transformer 모델인 Transcriptome Transformer (TxT)를 제안한다. TxT는 생존 예측을 주 과제로 설정하고, 임상 변수를 보조 과제로 함께 학습함으로써 전사체 패턴에 대한 표현 학습을 향상시킨다. 다중 헤드 attention 메커니즘을 통해 유전자–유전자 상호작용을 명시적으로 모델링하며, 전사체 특성을 반영한 새로운 positional embedding 기법을 도입하여 예측 성능을 향상시키는 동시에 모델의 해석 가능성을 유지한다. 다수의 암 코호트에 대한 검증을 통해 TxT가 임상 예후와 연관된 생물학적으로 의미 있는 유전자 기여를 효과적으로 포착할 수 있음을 확인하였다.

    둘째, 단일세포 전사체 약 9천5백만 건으로 사전 학습된 Transformer 모델인 Geneformer로부터 도출한 유전자 수준 임베딩을 조건으로 사용하는 분자 생성 모델 GGIFragGPT를 제안한다. 전사체 교란 시그니처로부터 생성된 유전자 임베딩을 기반으로 fragment 단위의 자기회귀적 분자 생성 과정을 수행함으로써, 화학적으로 타당하고 표현형과 연관된 분자를 안정적으로 생성할 수 있다. shRNA 유도 발현 시그니처를 활용한 CDK7 표적 분자 생성 사례를 통해, 제안한 모델이 목표 생물학적 표현형과 정합성이 높은 분자를 생성할 수 있음을 보였으며, attention 기반 분석을 통해 생성 과정의 해석 가능성 또한 제시하였다.

    종합적으로, 본 연구는 전사체 데이터를 임상 예측과 표현형 기반 분자 설계에 효과적으로 활용하기 위한 통합적 계산 프레임워크를 제시한다. Transformer 기반 모델에 유전자–유전자 상호작용, 임상 변수, 전사체 교란 정보를 명시적으로 통합함으로써, 데이터 기반 정밀의료 및 치료제 개발을 위한 해석 가능하고 생물학적으로 타당한 접근법을 제공한다.

    더보기

    목차 (Table of Contents)

    • Chapter 1 Introduction 1
    • 1.1 Background 4
    • 1.1.1 Transcriptomics and its role in biomedical research 4
    • 1.1.2 Transformer in computational biology 5
    • 1.1.3 Integrative multi-task modeling in biomedicine 6
    • Chapter 1 Introduction 1
    • 1.1 Background 4
    • 1.1.1 Transcriptomics and its role in biomedical research 4
    • 1.1.2 Transformer in computational biology 5
    • 1.1.3 Integrative multi-task modeling in biomedicine 6
    • 1.1.4 Transcriptome-guided phenotypic drug discovery 6
    • 1.1.5 Fragment-based molecular generation in drug discovery 8
    • 1.2 Research questions on transcriptome-based gene interactions 9
    • 1.2.1 Modeling gene?gene interactions for prognosis 9
    • 1.2.2 Generating therapeutic molecules using transcriptomic perturbations 11
    • 1.2.3 Toward interpretable transcriptome-informed modeling 11
    • 1.3 Outline of the thesis 12
    • Chapter 2 Improving patient survival prediction via multi-task learning of transcriptomic and clinical features 14
    • 2.1 Background 14
    • 2.1.1 Motivation 14
    • 2.1.2 Challenge 17
    • 2.1.3 Approach 18
    • 2.1.4 Related work 19
    • 2.2 Materials and Methods 20
    • 2.2.1 Formulating transcriptomic data for Transformer-based modeling 20
    • 2.2.2 Model architecture 23
    • 2.2.3 Pre-training gene embeddings on biological networks 23
    • 2.2.4 Transcriptome-based positional embedding 25
    • 2.2.5 Multi-task optimization technique 26
    • 2.2.6 Differential attention analysis 30
    • 2.2.7 Gene interaction modeling via multi-head attention 31
    • 2.2.8 Datasets 32
    • 2.2.9 Data preprocessing 37
    • 2.2.10 Model training 38
    • 2.3 Results 45
    • 2.3.1 Predictive performance across multiple datasets 45
    • 2.3.2 t-SNE visualization of patient embeddings 54
    • 2.3.3 Enhanced biological interpretability by clinical features 56
    • 2.3.4 Biological insights from attention-based gene interactions 57
    • 2.3.5 Model ablation and perturbation-based interpretation 60
    • 2.4 Discussion 71
    • Chapter 3 Transcriptome-conditioned molecule generation via gene interaction-aware fragment modeling with a GPT-based architecture 73
    • 3.1 Background 73
    • 3.1.1 Motivation 73
    • 3.1.2 Challenge 75
    • 3.1.3 Approach 76
    • 3.1.4 Related work 78
    • 3.2 Materials and Methods 80
    • 3.2.1 Computational formulation of transcriptome-to-molecule generation 80
    • 3.2.2 Dataset construction and preprocessing 83
    • 3.2.3 Gene embedding extraction using Geneformer 83
    • 3.2.4 Fragment-based molecular representation 84
    • 3.2.5 Determining fragment order 84
    • 3.2.6 Model architecture 85
    • 3.2.7 Training procedure 87
    • 3.2.8 Molecule generation strategy 87
    • 3.2.9 Gene-level interpretability via attention weights 89
    • 3.2.10 Target-oriented molecule generation guided by shRNA-induced transcriptomic signatures 90
    • 3.2.11 Structure-based selection of CDK7-targeting molecules via docking and interaction analysis 91
    • 3.3 Results 92
    • 3.3.1 Comparison of molecule generation performance 92
    • 3.3.2 Attention-based recovery of target genes 96
    • 3.3.3 Target-specific molecule generation guided by shRNA-induced transcriptomic profiles 99
    • 3.3.4 Molecule generation from disease transcriptomes 101
    • 3.3.5 CDK7-specific ligand discovery guided by transcriptomic perturbation and structure-based filtering 103
    • 3.4 Discussion 112
    • Chapter 4 Conclusion 115
    • Bibliography 119
    • 국문초록 138
    • 감사의 글 140
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼