단백질 서열 설계는 방대한 탐색 공간과 서열 변화의 민감성으로 인해 효율적인 최적화가 필요하다. 본 연구는 트랜스포머 인코더 기반 단백질 언어 모델과 생성 흐름 네트워크(GFlowNet)를 결...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
단백질 서열 설계는 방대한 탐색 공간과 서열 변화의 민감성으로 인해 효율적인 최적화가 필요하다. 본 연구는 트랜스포머 인코더 기반 단백질 언어 모델과 생성 흐름 네트워크(GFlowNet)를 결...
단백질 서열 설계는 방대한 탐색 공간과 서열 변화의 민감성으로 인해 효율적인 최적화가 필요하다. 본 연구는 트랜스포머 인코더 기반 단백질 언어 모델과 생성 흐름 네트워크(GFlowNet)를 결합하여, 특성 예측 보상을 활용한 새로운 서열 최적화 방식을 제안한다. 제안 모델은 위치별 아미노산 분포 예측, 변이 서열 생성, 특성 기반 보상 산출, GFlowNet 학습 단계로 구성 되며 보상에 비례한 확률로 다양한 고품질 변이를 탐색한다. 형광 단백질 CreiLOV를 대상으로 한 실험에서 제안된 모델은 원본 대비 1.2513% 향상된 로그 형광강도를 기록하며 PPO, GRPO 및 지도학습 기반 재학습 모델보다 우수한 성능을 보였다. 또한 구조적 유사성을 유지하면서 기능적 개선을 달성함을 Cα-RMSD 분석으로 확인하였다. 본 연구는 긴 단백질 서열에서도 안정적이고 다양한 탐색이 가능함을 보이며 자동화 단백질 설계의 유망한 방향성을 제시한다.
다국어 초록 (Multilingual Abstract)
Protein sequence design requires exploring vast combinatorial spaces in which small mutations can significantly alter structural and functional properties. To address the limitations of supervised fine-tuning and instability in reinforcement-learning...
Protein sequence design requires exploring vast combinatorial spaces in which small mutations can significantly alter structural and functional properties. To address the limitations of supervised fine-tuning and instability in reinforcement-learning–based optimization, we propose a GFlowNet-driven framework that integrates a transformer-based protein language model with sequence-property rewards. The method predicts amino-acid probability distributions, generates mutation candidates using threshold-based sampling, evaluates functional properties with a pretrained ensemble predictor, and trains a Sub-Trajectory Balance GFlowNet to sample sequences proportionally to their rewards. Using the oxygen-independent fluorescent protein CreiLOV and ESM-2 650M as the base model, the proposed approach achieves a 1.2513% increase in predicted log-fluorescence over the wild type, outperforming PPO-based fine-tuning and supervised fine-tuning. Structural analysis via Cα-RMSD confirms that the generated sequences maintain high similarity to the native fold while improving functional scores. These results demonstrate that GFlowNet-based optimization enables stable, diverse, and effective exploration of long protein sequences, offering a promising direction for automated protein design.
목차 (Table of Contents)