RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    번역 후 수정 부위 예측을 위한 딥러닝 접근법 : 변환기 기반 시퀀스 모델에서 즉시 조정되는 생성 아키텍처까지 = Deep Learning Approaches for Post-Translational Modification Site Prediction: From Transformer-Based Sequence Models to Prompt-Tuned Generative Architectures

    한글로보기

    https://www.riss.kr/link?id=T17370128

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    단백질의 번역 후 변형(Post-translational modifications, PTMs)은 단백질의 구조와 기능을 조절하는 중요한 요소로, 신호 전달, 단백질 안정성, 세포 내 위치 조절 등 다양한 생물학적 과정에서 중심적인 역할을 수행합니다. PTM 위치의 정확한 예측은 이러한 조절 메커니즘과 질병 연관 변화를 이해하는 데 필수적입니다. 하지만, 불균형한 데이터, 이질적인 모티프, 특히 저자원 또는 덜 연구된 PTM 유형에 대한 제한된 라벨 데이터로 인해 PTM 예측은 여전히 도전적인 과제로 남아 있습니다. 본 연구에서는 이러한 문제를 해결하기 위해 두 가지 상호보완적인 트랜스포머 기반 모델, DL-SPhos와 PTMGPT2를 제안하였습니다. 두 모델은 각각 판별 기반(discriminative) 및 생성 기반(generative) 학습 전략을 통해 PTM 예측의 새로운 방향을 제시합니다.

    DL-SPhos는 세린(Serine) 인산화 예측에 특화된 판별형 트랜스포머 모델입니다. 이 모델은 트랜스포머 인코더와 분류기(classification head)를 결합하여 문맥 기반의 잔기 수준 정보를 반영한 예측을 수행합니다. DL-SPhos는 어텐션 메커니즘을 통해 키나아제 특이적 모티프를 식별하며, 직관적인 해석이 가능합니다. 다양한 벤치마크에서 평가한 결과, DL-SPhos는 키나아제 비특이적 및 특이적 설정 모두에서 높은 정밀도를 기록하며 기존 분류 모델들을 능가하였습니다.
    반면, PTMGPT2는 PTM 예측을 위한 생성 기반 프레임워크를 제시합니다. 이 모델은 PTM 예측 문제를 프롬프트 조건 기반(sequence-to-modification) 생성 문제로 재정의하여, 예측하고자 하는 PTM 유형, 서열 문맥, 또는 생물학적 제약조건 등을 포함한 구조화된 프롬프트를 입력으로 받아 유연하게 잔기–변형 쌍을 출력합니다. 프롬프트 튜닝과 토큰 설계(token engineering)를 통해 예측 작업을 세밀하게 조절할 수 있으며, 어텐션 분석은 서열–라벨 간의 생물학적으로 해석 가능한 관계를 제공합니다.
    총 19종의 PTM 유형을 대상으로 한 벤치마크 평가에서, PTMGPT2는 MCC, F1-score, Precision 등 핵심 지표 전반에서 최신 기법들을 안정적으로 능가하였습니다. 사례 기반 생물학적 분석에서도 모델의 유효성이 입증되었는데, 어텐션 헤드를 통한 분석에서는 다양한 키나아제 계열에서 알려진 PTM 모티프들이 성공적으로 복원되었으며, 생성적 돌연변이 시뮬레이션을 통해 암 관련 유전자들(TP53, BRAF, RAF1) 근처의 변형 민감 영역도 효과적으로 식별하였습니다. 이러한 결과는 PTMGPT2가 단순한 예측을 넘어서 기능적 서열 요소에 대한 기전적 통찰을 제공함을 보여줍니다.
    결론적으로, DL-SPhos와 PTMGPT2는 해석 가능성, 일반화 능력, 생물학적 타당성을 결합하여 PTM 예측의 수준을 한층 끌어올렸습니다. DL-SPhos는 특정 키나아제에 기반한 정밀 예측에 적합하며, PTMGPT2는 다양한 PTM 유형을 아우를 수 있는 프롬프트 기반의 확장 가능한 플랫폼을 제공합니다. 두 모델의 통합은 향후 변이 해석, 단백질체학, 정밀 의학 등의 분야에서 활용될 수 있는 단백질 모델링 파이프라인의 새로운 가능성을 제시합니다.
    번역하기

    단백질의 번역 후 변형(Post-translational modifications, PTMs)은 단백질의 구조와 기능을 조절하는 중요한 요소로, 신호 전달, 단백질 안정성, 세포 내 위치 조절 등 다양한 생물학적 과정에서 중심적...

    단백질의 번역 후 변형(Post-translational modifications, PTMs)은 단백질의 구조와 기능을 조절하는 중요한 요소로, 신호 전달, 단백질 안정성, 세포 내 위치 조절 등 다양한 생물학적 과정에서 중심적인 역할을 수행합니다. PTM 위치의 정확한 예측은 이러한 조절 메커니즘과 질병 연관 변화를 이해하는 데 필수적입니다. 하지만, 불균형한 데이터, 이질적인 모티프, 특히 저자원 또는 덜 연구된 PTM 유형에 대한 제한된 라벨 데이터로 인해 PTM 예측은 여전히 도전적인 과제로 남아 있습니다. 본 연구에서는 이러한 문제를 해결하기 위해 두 가지 상호보완적인 트랜스포머 기반 모델, DL-SPhos와 PTMGPT2를 제안하였습니다. 두 모델은 각각 판별 기반(discriminative) 및 생성 기반(generative) 학습 전략을 통해 PTM 예측의 새로운 방향을 제시합니다.

    DL-SPhos는 세린(Serine) 인산화 예측에 특화된 판별형 트랜스포머 모델입니다. 이 모델은 트랜스포머 인코더와 분류기(classification head)를 결합하여 문맥 기반의 잔기 수준 정보를 반영한 예측을 수행합니다. DL-SPhos는 어텐션 메커니즘을 통해 키나아제 특이적 모티프를 식별하며, 직관적인 해석이 가능합니다. 다양한 벤치마크에서 평가한 결과, DL-SPhos는 키나아제 비특이적 및 특이적 설정 모두에서 높은 정밀도를 기록하며 기존 분류 모델들을 능가하였습니다.
    반면, PTMGPT2는 PTM 예측을 위한 생성 기반 프레임워크를 제시합니다. 이 모델은 PTM 예측 문제를 프롬프트 조건 기반(sequence-to-modification) 생성 문제로 재정의하여, 예측하고자 하는 PTM 유형, 서열 문맥, 또는 생물학적 제약조건 등을 포함한 구조화된 프롬프트를 입력으로 받아 유연하게 잔기–변형 쌍을 출력합니다. 프롬프트 튜닝과 토큰 설계(token engineering)를 통해 예측 작업을 세밀하게 조절할 수 있으며, 어텐션 분석은 서열–라벨 간의 생물학적으로 해석 가능한 관계를 제공합니다.
    총 19종의 PTM 유형을 대상으로 한 벤치마크 평가에서, PTMGPT2는 MCC, F1-score, Precision 등 핵심 지표 전반에서 최신 기법들을 안정적으로 능가하였습니다. 사례 기반 생물학적 분석에서도 모델의 유효성이 입증되었는데, 어텐션 헤드를 통한 분석에서는 다양한 키나아제 계열에서 알려진 PTM 모티프들이 성공적으로 복원되었으며, 생성적 돌연변이 시뮬레이션을 통해 암 관련 유전자들(TP53, BRAF, RAF1) 근처의 변형 민감 영역도 효과적으로 식별하였습니다. 이러한 결과는 PTMGPT2가 단순한 예측을 넘어서 기능적 서열 요소에 대한 기전적 통찰을 제공함을 보여줍니다.
    결론적으로, DL-SPhos와 PTMGPT2는 해석 가능성, 일반화 능력, 생물학적 타당성을 결합하여 PTM 예측의 수준을 한층 끌어올렸습니다. DL-SPhos는 특정 키나아제에 기반한 정밀 예측에 적합하며, PTMGPT2는 다양한 PTM 유형을 아우를 수 있는 프롬프트 기반의 확장 가능한 플랫폼을 제공합니다. 두 모델의 통합은 향후 변이 해석, 단백질체학, 정밀 의학 등의 분야에서 활용될 수 있는 단백질 모델링 파이프라인의 새로운 가능성을 제시합니다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Post-translational modifications (PTMs) are critical modulators of protein structure and function, playing central roles in signal transduction, protein stability, and cellular localization. The accurate prediction of PTM sites is essential for understanding regulatory mechanisms and disease-associated alterations. However, computational PTM prediction is hindered by imbalanced data, heterogeneous motifs, and limited labeled datasets, especially for low-resource or less-studied PTM types. This work addresses these challenges by introducing two transformer-based models: DL-SPhos and PTMGPT2 that offer complementary approaches to PTM site prediction through discriminative and generative learning strategies.
    DL-SPhos is a discriminative transformer model designed specifically for serine phosphorylation prediction. It combines a transformer encoder with a classification head, enabling residue-level scoring informed by contextual embeddings. The model leverages attention mechanisms to identify kinase-specific motifs and offers interpretability. Evaluated across multiple benchmarks, DL-SPhos demonstrates high precision in serine phosphorylation tasks, outperforming traditional classifiers in both kinase-agnostic and kinase-specific settings.

    In contrast, PTMGPT2 introduces a generative modeling framework for PTM prediction. It frames PTM site prediction as a prompt-conditioned sequence generation task, accepting structured prompts that define target PTM types, sequence context, or biological constraints. This enables the model to output residue–modification labels in a flexible and scalable manner without retraining. Prompt tuning and token engineering support fine-grained control over prediction tasks, while attention heads offer biologically coherent insights into sequence–label dependencies.
    Benchmark evaluations across 19 PTM types reveal that PTMGPT2 consistently outperforms state-of-the-art models on key metrics including MCC, F1 score, and precision. In-depth case studies further validate PTMGPT2’s biological utility: attention head analysis recovers known PTM motifs across kinase families, and generative perturbation experiments reveal mutation-sensitive hotspots near phosphosites in cancer-associated genes such as TP53, BRAF, and RAF1. These results demonstrate that PTMGPT2 not only supports accurate PTM prediction but also facilitates mechanistic insight into functional sequence elements.
    Together, DL-SPhos and PTMGPT2 advance the state of PTM site prediction by combining interpretability, generalization, and biological relevance. DL-SPhos provides a robust solution for targeted kinase-specific tasks, while PTMGPT2 offers a prompt-tuned, extensible platform for large-scale PTM tasks. Their integration into protein modeling pipelines holds promise for downstream applications in variant interpretation, proteomics, and precision medicine.
    번역하기

    Post-translational modifications (PTMs) are critical modulators of protein structure and function, playing central roles in signal transduction, protein stability, and cellular localization. The accurate prediction of PTM sites is essential for unders...

    Post-translational modifications (PTMs) are critical modulators of protein structure and function, playing central roles in signal transduction, protein stability, and cellular localization. The accurate prediction of PTM sites is essential for understanding regulatory mechanisms and disease-associated alterations. However, computational PTM prediction is hindered by imbalanced data, heterogeneous motifs, and limited labeled datasets, especially for low-resource or less-studied PTM types. This work addresses these challenges by introducing two transformer-based models: DL-SPhos and PTMGPT2 that offer complementary approaches to PTM site prediction through discriminative and generative learning strategies.
    DL-SPhos is a discriminative transformer model designed specifically for serine phosphorylation prediction. It combines a transformer encoder with a classification head, enabling residue-level scoring informed by contextual embeddings. The model leverages attention mechanisms to identify kinase-specific motifs and offers interpretability. Evaluated across multiple benchmarks, DL-SPhos demonstrates high precision in serine phosphorylation tasks, outperforming traditional classifiers in both kinase-agnostic and kinase-specific settings.

    In contrast, PTMGPT2 introduces a generative modeling framework for PTM prediction. It frames PTM site prediction as a prompt-conditioned sequence generation task, accepting structured prompts that define target PTM types, sequence context, or biological constraints. This enables the model to output residue–modification labels in a flexible and scalable manner without retraining. Prompt tuning and token engineering support fine-grained control over prediction tasks, while attention heads offer biologically coherent insights into sequence–label dependencies.
    Benchmark evaluations across 19 PTM types reveal that PTMGPT2 consistently outperforms state-of-the-art models on key metrics including MCC, F1 score, and precision. In-depth case studies further validate PTMGPT2’s biological utility: attention head analysis recovers known PTM motifs across kinase families, and generative perturbation experiments reveal mutation-sensitive hotspots near phosphosites in cancer-associated genes such as TP53, BRAF, and RAF1. These results demonstrate that PTMGPT2 not only supports accurate PTM prediction but also facilitates mechanistic insight into functional sequence elements.
    Together, DL-SPhos and PTMGPT2 advance the state of PTM site prediction by combining interpretability, generalization, and biological relevance. DL-SPhos provides a robust solution for targeted kinase-specific tasks, while PTMGPT2 offers a prompt-tuned, extensible platform for large-scale PTM tasks. Their integration into protein modeling pipelines holds promise for downstream applications in variant interpretation, proteomics, and precision medicine.

    더보기

    목차 (Table of Contents)

    • Table of Contents i
    • List of Figure vi
    • List of Tables xv
    • Abstract xvii
    • 요약 xix
    • Table of Contents i
    • List of Figure vi
    • List of Tables xv
    • Abstract xvii
    • 요약 xix
    • Chapter 1 Introduction 1
    • 1.1 Background 1
    • 1.2 Types of PTMs and Their Functional Roles in Cellular Processes 4
    • 1.3 Challenges in PTM Site Prediction 6
    • 1.4 From Classical Machine Learning to Deep Learning Approaches 7
    • 1.5 The Rise of Transformer-Based Models in Bioinformatics 9
    • 1.6 Research Objectives and Scope 10
    • 1.7 Conclusion 12
    • Chapter 2 Biological Foundation of PTMs 14
    • 2.1 Protein Sequence and Functional Domains 14
    • 2.2 Enzymes Involved in PTMs: Kinases, Transferases, and Ligases 15
    • 2.2.1 Kinases 16
    • 2.2.2 Transferases 16
    • 2.2.3 Ligases 17
    • 2.3 Major PTMs: Phosphorylation, Methylation, Acetylation, Glycosylation 18
    • 2.3.1 Phosphorylation 18
    • 2.3.2 Methylation 19
    • 2.3.3 Acetylation 19
    • 2.3.4 Glycosylation 20
    • 2.4 Serine Phosphorylation: Prevalence and Disease Link 21
    • 2.5 Biological Role of PTM Motifs in Mammalian Diseases 23
    • 2.5.1 Conserved Motifs as Regulatory Elements 23
    • 2.5.2 Disease Mechanisms via Motif Dysregulation 24
    • 2.5.3 Implications for PTM Prediction and Therapeutics 24
    • 2.6 Databases for PTM Annotation: UniProt, dbPTM, PTMD 25
    • 2.6.1 UniProtKB 26
    • 2.6.2 dbPTM 26
    • 2.6.3 PTMD 27
    • 2.7 Case Study: Serine Phosphorylation in Neurodegeneration and Cancer 28
    • 2.7.1 Neurodegeneration: Hyperphosphorylation of Tau Protein 28
    • 2.7.2 Cancer: Kinase Dysregulation and Serine Phosphorylation Cascades 29
    • 2.7.3 Translational Implications and Modeling Perspective 29
    • 2.8 Conclusion 30
    • Chapter 3 Transformer-Based Sequence Models for PTM Site Prediction 31
    • 3.1 Introduction to Transformer Architecture 31
    • 3.1.1 Key Components of the Transformer 31
    • 3.1.2 Advantages for PTM Prediction 32
    • 3.1.3 Application in This Thesis 33
    • 3.2 Limitations of Prior PTM Prediction Approaches 33
    • 3.2.1 Dependence on Handcrafted Features and Motif Libraries 34
    • 3.2.2 Poor Generalization Across Kinase Types 34
    • 3.2.3 Limited Interpretability 34
    • 3.3 Design and Training of DL-SPhos 35
    • 3.3.1 Model Architecture 36
    • 3.3.2 Dataset Construction 38
    • 3.3.3 Training Strategy 42
    • 3.3.4 Advantages of DL-SPhos Design 44
    • 3.4 Integration of Explainable AI (LIME) 45
    • 3.4.1 Motivation for Interpretability in PTM Prediction 45
    • 3.4.2 LIME Workflow for Protein Sequences 46
    • 3.4.3 Limitations and Future Extensions 46
    • 3.5 Evaluation on Benchmark and Independent Datasets 47
    • 3.5.1 Evaluation Metrics 47
    • 3.5.2 5-Fold Cross Validation Results 48
    • 3.5.3 Independent dataset results 49
    • 3.5.4 Kinase-specific benchmark dataset results 51
    • 3.6 Kinase-Specific Motif Discovery 63
    • 3.6.1 Protein Kinase–Based Motif Generation 63
    • 3.6.2 Analysis of Kinase-Specific Motifs Identified by DL-SPhos 64
    • 3.7 Case Study: Serine Phosphorylation in Mammalian Diseases 66
    • 3.8 Summary of DL-SPhos Contributions 67
    • 3.9 Conclusion 69
    • Chapter 4 Prompt-Tuned Generative Transformer for PTM Prediction 70
    • 4.1 Motivation for Prompt-Tuned Learning 70
    • 4.2 Overview of the PTMGPT2 Architecture 70
    • 4.3 Role of PROTGPT2 as a pretrained model 73
    • 4.3.1 Architecture Overview of ProtGPT2 73
    • 4.3.2 Protein-Centric Pretraining 73
    • 4.3.3 Advantages Over General NLP Backbones 74
    • 4.4 Dataset Preparation 74
    • 4.5 Custom Token Design and Prompt Structure 77
    • 4.5.1 Motivation for Prompt-Based Conditioning 79
    • 4.5.2 Custom Token Vocabulary 80
    • 4.5.3 Prompt Template Design 81
    • 4.6 Benchmark Evaluation 81
    • 4.7 Species-specific performance 85
    • 4.7.1 Methylation (R) 87
    • 4.7.2 Methylation (K) 87
    • 4.7.3 Acetylation (K) 87
    • 4.7.4 Formylation (K) 87
    • 4.7.5 Glutarylation (K) 88
    • 4.7.6 Glutathionylation (K) 88
    • 4.7.7 Hydroxylation (K) 88
    • 4.7.8 Hydroxylation (P) 88
    • 4.7.9 Malonylation (K) 89
    • 4.7.10 N-linked Glycosylation (N) 89
    • 4.7.11 O-linked Glycosylation (S,T) 89
    • 4.7.12 Phosphorylation (S,T) 89
    • 4.7.13 Phosphorylation (Y) 90
    • 4.7.14 S-Nitrosylation (C) 90
    • 4.7.15 S-Palmitoylation (C) 90
    • 4.7.16 Succinylation (K) 90
    • 4.7.17 Sumoylation (K) 91
    • 4.7.18 Ubiquitination (K) 91
    • 4.8 PTMGPT2 performance on different CD-HIT similarity cutoffs 91
    • 4.9 PTMGPT2 training and inference prompt designs 93
    • 4.10 Sequence–Label Dependency via Attention Heads 94
    • 4.10.1 Interpretable Attention Framework in PTMGPT2 94
    • 4.10.2 Algorithm – PSPM Construction from Attention Profiles 96
    • 4.10.3 Kinase Family-Specific Motif Insights 97
    • 4.10.4 PSPM Construction and Visualization 99
    • 4.11 PTMGPT2’s Mutation Hotspot Identification 100
    • 4.11.1 Background and Motivation 100
    • 4.11.2 Methodological Framework for Hotspot Detection 100
    • 4.11.3 Gene-Wise Results and Biological Interpretation 101
    • 4.11.4 Quantitative Evaluation and Visualization 102
    • 4.12 PTMGPT2 performance on recently deposited proteins 112
    • 4.12.1 Phosphorylation (T) 113
    • 4.12.2 Phosphorylation (Y) 113
    • 4.12.3 Phosphorylation (S) 114
    • 4.12.4 Methylation (K) 115
    • 4.12.5 Acetylation (K) 115
    • 4.13 Summary of PTMGPT2 Contributions 116
    • 4.13.1 Prompt-Based Generative Reformulation of PTM Prediction 116
    • 4.13.2 Backbone Adaptation via Protein Language Modeling 117
    • 4.13.3 Custom Prompt Engineering and Token Schema 117
    • 4.13.4 Interpretability via Attention and Mutation Sensitivity 118
    • 4.13.5 Extensible Foundation for Future Multi-Omics Applications 118
    • Chapter 5 Discussion 120
    • Chapter 6 Conclusion 124
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼