RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    서열 특성 보상을 활용한 생성 흐름 네트워크 기반 단백질 언어 모델의 최적화 = Optimization of Protein Language Models via GFlowNet with Sequence Property Rewards

    한글로보기

    https://www.riss.kr/link?id=T17380482

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    단백질 서열 설계는 방대한 탐색 공간과 서열 변화의 민감성으로 인해 효율적인 최적화가 필요하다. 본 연구는 트랜스포머 인코더 기반 단백질 언어 모델과 생성 흐름 네트워크(GFlowNet)를 결합하여, 특성 예측 보상을 활용한 새로운 서열 최적화 방식을 제안한다. 제안 모델은 위치별 아미노산 분포 예측, 변이 서열 생성, 특성 기반 보상 산출, GFlowNet 학습 단계로 구성 되며 보상에 비례한 확률로 다양한 고품질 변이를 탐색한다. 형광 단백질 CreiLOV를 대상으로 한 실험에서 제안된 모델은 원본 대비 1.2513% 향상된 로그 형광강도를 기록하며 PPO, GRPO 및 지도학습 기반 재학습 모델보다 우수한 성능을 보였다. 또한 구조적 유사성을 유지하면서 기능적 개선을 달성함을 Cα-RMSD 분석으로 확인하였다. 본 연구는 긴 단백질 서열에서도 안정적이고 다양한 탐색이 가능함을 보이며 자동화 단백질 설계의 유망한 방향성을 제시한다.
    번역하기

    단백질 서열 설계는 방대한 탐색 공간과 서열 변화의 민감성으로 인해 효율적인 최적화가 필요하다. 본 연구는 트랜스포머 인코더 기반 단백질 언어 모델과 생성 흐름 네트워크(GFlowNet)를 결...

    단백질 서열 설계는 방대한 탐색 공간과 서열 변화의 민감성으로 인해 효율적인 최적화가 필요하다. 본 연구는 트랜스포머 인코더 기반 단백질 언어 모델과 생성 흐름 네트워크(GFlowNet)를 결합하여, 특성 예측 보상을 활용한 새로운 서열 최적화 방식을 제안한다. 제안 모델은 위치별 아미노산 분포 예측, 변이 서열 생성, 특성 기반 보상 산출, GFlowNet 학습 단계로 구성 되며 보상에 비례한 확률로 다양한 고품질 변이를 탐색한다. 형광 단백질 CreiLOV를 대상으로 한 실험에서 제안된 모델은 원본 대비 1.2513% 향상된 로그 형광강도를 기록하며 PPO, GRPO 및 지도학습 기반 재학습 모델보다 우수한 성능을 보였다. 또한 구조적 유사성을 유지하면서 기능적 개선을 달성함을 Cα-RMSD 분석으로 확인하였다. 본 연구는 긴 단백질 서열에서도 안정적이고 다양한 탐색이 가능함을 보이며 자동화 단백질 설계의 유망한 방향성을 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Protein sequence design requires exploring vast combinatorial spaces in which small mutations can significantly alter structural and functional properties. To address the limitations of supervised fine-tuning and instability in reinforcement-learning–based optimization, we propose a GFlowNet-driven framework that integrates a transformer-based protein language model with sequence-property rewards. The method predicts amino-acid probability distributions, generates mutation candidates using threshold-based sampling, evaluates functional properties with a pretrained ensemble predictor, and trains a Sub-Trajectory Balance GFlowNet to sample sequences proportionally to their rewards. Using the oxygen-independent fluorescent protein CreiLOV and ESM-2 650M as the base model, the proposed approach achieves a 1.2513% increase in predicted log-fluorescence over the wild type, outperforming PPO-based fine-tuning and supervised fine-tuning. Structural analysis via Cα-RMSD confirms that the generated sequences maintain high similarity to the native fold while improving functional scores. These results demonstrate that GFlowNet-based optimization enables stable, diverse, and effective exploration of long protein sequences, offering a promising direction for automated protein design.
    번역하기

    Protein sequence design requires exploring vast combinatorial spaces in which small mutations can significantly alter structural and functional properties. To address the limitations of supervised fine-tuning and instability in reinforcement-learning...

    Protein sequence design requires exploring vast combinatorial spaces in which small mutations can significantly alter structural and functional properties. To address the limitations of supervised fine-tuning and instability in reinforcement-learning–based optimization, we propose a GFlowNet-driven framework that integrates a transformer-based protein language model with sequence-property rewards. The method predicts amino-acid probability distributions, generates mutation candidates using threshold-based sampling, evaluates functional properties with a pretrained ensemble predictor, and trains a Sub-Trajectory Balance GFlowNet to sample sequences proportionally to their rewards. Using the oxygen-independent fluorescent protein CreiLOV and ESM-2 650M as the base model, the proposed approach achieves a 1.2513% increase in predicted log-fluorescence over the wild type, outperforming PPO-based fine-tuning and supervised fine-tuning. Structural analysis via Cα-RMSD confirms that the generated sequences maintain high similarity to the native fold while improving functional scores. These results demonstrate that GFlowNet-based optimization enables stable, diverse, and effective exploration of long protein sequences, offering a promising direction for automated protein design.

    더보기

    목차 (Table of Contents)

    • 제1장 서론 1
    • 제2장 배경지식 및 관련연구 4
    • 제1절 배경지식 4
    • 1. 강화학습 4
    • 2. 생성 흐름 네트워크 6
    • 제1장 서론 1
    • 제2장 배경지식 및 관련연구 4
    • 제1절 배경지식 4
    • 1. 강화학습 4
    • 2. 생성 흐름 네트워크 6
    • 제2절 관련연구 8
    • 1. 단백질 서열 생성 8
    • 2. 단백질 특성 예측 9
    • 3. 단백질 서열 설계 자동화 9
    • 4. GFlowNet 기반 LLM 최적화 10
    • 제3장 제안 모델 12
    • 제1절 제안 모델의 구조 12
    • 1. Amino Acid Probability Prediction 13
    • 2. Mutated Sequence Generation 14
    • 3. Sequence Feature Prediction 16
    • 4. GFlowNet Tuning 17
    • 제4장 실험 및 성능 평가 19
    • 제1절 실험 환경 및 데이터 19
    • 1. 대상 단백질 및 데이터셋 19
    • 2. 비교 모델 20
    • 제2절 실험 결과 분석 22
    • 제5장 결론 및 향후 계획 31
    • 제1절 제안 모델의 기여 31
    • 제2절 한계 및 향후 계획 31
    • 참고문헌 33
    • 부록 38
    • ABSTRACT 42
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼