병원성 변이의 신뢰성 높은 예측은 유전체 정보를 활용한 정확한 진단과 맞춤형 치료를 목표로 하는 정밀 의학에서 매우 중요한 역할을 한다. 그러나 이러한 변이들의 실험적 검증은 비용이...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17315108
서울 : 서울대학교 대학원, 2025
학위논문(박사) -- 서울대학교 대학원 , 협동과정생물정보학전공 , 2025. 8
2025
영어
574.8732
서울
xi, 158 ; 26 cm
지도교수: 김주한
I804:11032-000000191519
0
상세조회0
다운로드병원성 변이의 신뢰성 높은 예측은 유전체 정보를 활용한 정확한 진단과 맞춤형 치료를 목표로 하는 정밀 의학에서 매우 중요한 역할을 한다. 그러나 이러한 변이들의 실험적 검증은 비용이...
병원성 변이의 신뢰성 높은 예측은 유전체 정보를 활용한 정확한 진단과 맞춤형 치료를 목표로 하는 정밀 의학에서 매우 중요한 역할을 한다. 그러나 이러한 변이들의 실험적 검증은 비용이 많이 들고 시간 소모가 크기 때문에 대규모 연구에서는 현실적으로 어렵다. 이러한 한계를 극복하기 위해 유전 변이의 잠재적 영향을 예측하는 다양한 계산적 방법들이 개발되었다. 본 연구에서는 희귀 비동의 단일염기변이(nsSNV)에 대한 기존 예측 방법들의 성능을 체계적으로 평가하고, 이들의 특성과 한계를 조명한다. 또한 이러한 한계를 극복하여 예측 성능을 향상시킨 새로운 방법인 PRP(병원성 위험 예측)를 소개한다.
본 논문은 총 4개의 장으로 구성되어 있다. 1장에서는 관련 문헌 조사를 통해 병원성 변이 예측에 관한 선행 연구들을 설명하고, 본 연구의 목적을 제시한다.
2장에서는 기존 병원성 예측 방법들의 특성을 요약하고, 그 성능을 비교 분석한다. 지금까지 다양한 예측 방법들의 성능을 비교한 연구가 있었지만, 희귀 변이에 대한 이 방법들의 성능은 아직 체계적으로 비교 평가되지 않았다. 본 연구에서는 최신 ClinVar 데이터셋을 활용하여 총 28개의 병원성 예측 방법들의 성능을 비교 평가하였으며, 특히 희귀 변이와 다양한 대립 유전자 빈도 (AF) 범위에 초점을 맞추어 성능을 분석하였다. 대부분의 방법들은 비동의 단일염기변이 중 미스센스와 start_lost 변이만을 다뤘으며, 데이터셋 내 변이들 중 약 10% 정도의 예측 점수 결측률이 관찰되었다. 보존성 정보, 다른 예측 점수, 대립 유전자 빈도를 특성으로 포함한 MetaRNN과 ClinPred가 희귀 변이에 대해 가장 높은 예측 성능을 보였다. 대부분 방법들에서 특이도가 민감도보다 낮았다. 다양한 대립 유전자 빈도 범위에서 대립 유전자 빈도가 감소함에 따라 대부분의 성능 지표가 전반적으로 하락하는 경향을 보였으며, 특히 특이도의 감소가 두드러졌다. 이러한 결과는 희귀 변이의 병원성 예측에 있어 각 방법의 강점과 한계를 보여주며, 향후 예측 모델의 개선 방향을 제시한다.
3장에서는 희귀 비동의 단일염기 변이에 대한 병원성 위험을 예측을 위한 새로운 방법인 PRP를 제시한다. PRP는 빈도, 보존도, 치환 지표, 유전자 내성의 네가지 범주에 걸친 총 34개의 특성들을 활용하여 견고한 성능과 해석 가능한 예측을 목표로 설계되었다. 최적의 모델을 선정을 위해 다섯 가지 머신러닝(ML) 알고리즘을 비교하였다. 하이퍼파라미터 최적화는 Optuna를 사용하여 수행하였으며, SHAP(Shapley Additive exPlanations)를 통해 특성 중요도를 분석하였다. PRP는 ClinVar 데이터를 학습에 사용하였고, 세 개의 독립적인 테스트 데이터셋을 통해 성능을 평가하였으며. 20개의 다른 예측 방법들과 성능을 비교하였다. PRP는 8가지 성능 지표 전반에 걸쳐 일관되게 최상위 성능을 나타냈다. 특히 병원성 변이를 과대평가하지 않으면서도 높은 민감도와 특이도를 동시에 달성하였으며, 일반 변이뿐만 아니라 희귀 변이까지 아우르는 변이 예측에서 견고함을 입증하였다.
4장에서는 본 연구의 주요 결과를 요약하고, 향후 연구 및 임상 적용에 대한 개선점을 논의하며, 희귀 비동의 변이의 병원성 예측을 더욱 개선하기 위한 잠재적 방향을 제시한다.
다국어 초록 (Multilingual Abstract)
Reliable prediction of pathogenic variants plays a crucial role in precision medicine, which aims to provide accurate diagnoses and personalized treatments through genomic information. However, experimental validation of these variants is often imprac...
Reliable prediction of pathogenic variants plays a crucial role in precision medicine, which aims to provide accurate diagnoses and personalized treatments through genomic information. However, experimental validation of these variants is often impractical in large-scale studies due to high costs and time-consuming. To overcome these limitations, numerous computational methods have been developed to predict the potential impact of genetic variants. This study systematically evaluates the performance of existing prediction methods in rare nonsynonymous single nucleotide variants (nsSNVs), highlights their characteristics and limitations. Furthermore, it presents a novel method, PRP (Pathogenic Risk Prediction), which improves prediction performance by overcoming these limitations.
This thesis consists of four chapters. Chapter 1 provides an overview of previous studies on pathogenic variant prediction through a review of the relevant literature and states the purpose of this study.
Chapter 2 provides a summary of the characteristics and a comparative analysis of the performance of existing pathogenicity prediction methods. Although previous studies have compared the performance of various prediction methods, the performance of these methods specifically on rare variants has not yet been systematically assessed. In this study, the performance of 28 pathogenicity prediction methods was assessed using the latest ClinVar dataset, with a focus on rare variants and various allele frequency (AF) ranges. Most methods focused on missense and start-lost variants, covering only a subset of nonsynonymous SNVs. The average missing rate of approximately 10% was observed among the variants in the dataset, indicating that prediction scores were unavailable for them. MetaRNN and ClinPred, which incorporated conservation, other prediction scores, and AFs as features, demonstrated the highest predictive power on rare variants. For most methods, specificity was lower than sensitivity. Across various AF ranges, most performance metrics tended to decline as AF decreased, with specificity showing a particularly large decline. These results provide insights into the strengths and limitations of each method in predicting the pathogenic of rare variants, which may guide future improvements in predictive models.
Chapter3 presents a novel method, PRP, for pathogenic risk prediction of rare nsSNVs. PRP was designed to provide robust performance and interpretable predictions using thirty-four features across four categories: frequency, conservation score, substitution metrics, and gene intolerance. Five machine-learning (ML) algorithms were compared to select the optimal model. Hyperparameter optimization was conducted using Optuna, and feature importance was analyzed using Shapley Additive exPlanations (SHAP). PRP used ClinVar data for training and evaluated performance using three independent test datasets and compared it with that of 20 other prediction methods. PRP consistently outperformed state-of-the-art methods across all eight performance metrics. In addition to achieving high sensitivity and high specificity without overestimating the number of pathogenic variants, PRP demonstrates robustness in predicting variants across the spectrum, including both common and rare variants.
Chapter 4 summarizes the main findings of this study, discusses areas for improvement in future research and clinical applications, and suggests potential directions for further enhancing the pathogenic prediction of variants.
목차 (Table of Contents)
[제18회 김옥길기념강좌] 인공지능, 감정, 휴머니즘(Human-Compatible Artificial Intelligence’)’
이화여자대학교 스튜어드 러셀Machine Learning for Data Science
K-MOOC 고려대학교 정태수, 주재걸, 석준희누구나 할 수 있는 AI 머신 러닝[AI Machine Learning Zero To All]
K-MOOC 인하공업전문대학 이세훈누구나 할 수 있는 데이터 분석과 인공지능[Data Analysis and Artificial Intelligence for Everyone]
K-MOOC 인하공업전문대학 이세훈누구나 할 수 있는 데이터 분석과 인공지능[Data Analysis and Artificial Intelligence for Everyone]
K-MOOC 인하공업전문대학 이세훈