본 연구에서는 대규모 유전체-표현형 연계 데이터인 UK Biobank (UKB)를 활용하여, 전장엑솜서열(Whole Exome Sequencing, WES) 기반 유전 변이와 다양한 임상 표현형 간의 관계를 포괄적으로 분석할 수 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17315332
서울 : 서울대학교 대학원, 2025
학위논문(박사) -- 서울대학교 대학원 , 협동과정생물정보학전공 , 2025. 8
2025
한국어
UK Biobank ; PheWAS ; 유전 변이 ; 기능 해석 ; 웹 기반 도구
574.8732
서울
53 ; 26 cm
지도교수: 김주한
I804:11032-000000192913
0
상세조회0
다운로드본 연구에서는 대규모 유전체-표현형 연계 데이터인 UK Biobank (UKB)를 활용하여, 전장엑솜서열(Whole Exome Sequencing, WES) 기반 유전 변이와 다양한 임상 표현형 간의 관계를 포괄적으로 분석할 수 ...
본 연구에서는 대규모 유전체-표현형 연계 데이터인 UK Biobank (UKB)를 활용하여, 전장엑솜서열(Whole Exome Sequencing, WES) 기반 유전 변이와 다양한 임상 표현형 간의 관계를 포괄적으로 분석할 수 있는 바이오마커 군집 기반의 새로운 접근법을 고안하였다. 또한, 이러한 분석 수행과 결과의 시각화를 지원하는 웹 기반 분석 도구인 UK-IGV (UK Biobank–Based Interpretation of Genetic Variants)를 개발하였다.
약 20만 명의 WES 변이 데이터와 임상 데이터를 통해 65개 임상 검사 항목을 표현형으로 변이 별 Phenome-Wide Association Study (PheWAS) 분석을 수행하였다. 전체 집단뿐만 아니라 연령대 별, 성별, 질병 집단 등 다양한 하위 집단에 대한 분석 결과도 함께 제공하여, 유전 변이가 특정 집단에서 표현형에 미치는 영향을 확인할 수 있도록 하였다. 통계적으로 유의한 PheWAS 결과는 바이올린 플롯(Violin plot)과 맨하탄 플롯(Manhattan plot) 시각화를 통해 해석의 직관성을 높였다.
또한, 외부 주석 데이터를 이용하여 변이 기능 예측 점수(SIFT, PolyPhen-2, CADD, PrimateAI, AlphaMissense), 질환 관련 정보(ClinVar, OMIM), 인종 별 변이 빈도(1KGP, gnomAD, UKB)까지 함께 제공하여, 변이의 기능에 대한 해석을 지원하였다.
아울러, PheWAS 결과 바탕으로 변이의 feature 벡터 및 표현형과 군집에 대한 영향력 지표를 정의하여 기계학습 기반 군집화를 수행하였다. 이로써, 변이 간 기능적 유사성을 반영한 군집 구조를 구축하여, 각 군집에 대해 우세한 바이오마커(Dominant Biomarker)를 식별하고 기능적 패턴을 효과적으로 파악할 수 있도록 하였다. 더 나아가, Dominant Biomarker 군집 단위의 enrichment 분석을 통해, 군집 별 생물학적 기전에 대한 통찰력을 제공할 수 있도록 설계하였다.
본 도구는 유전 변이의 표현형 연관성, 시각화, 외부 주석 정보, 군집 기반 비교 해석을 통합한 분석 플랫폼으로, 바이오뱅크 기반 유전체 연구에서 후보 유전 변이 및 유전자 발굴과 검증에 유용하게 활용될 수 있을 것으로 기대된다.
다국어 초록 (Multilingual Abstract)
In this study, we propose a novel biomarker cluster-based approach for comprehensively analyzing the relationship between genetic variants derived from Whole Exome Sequencing (WES) and a wide range of clinical phenotypes, utilizing the large-scale gen...
In this study, we propose a novel biomarker cluster-based approach for comprehensively analyzing the relationship between genetic variants derived from Whole Exome Sequencing (WES) and a wide range of clinical phenotypes, utilizing the large-scale genomic-phenotypic resource of the UK Biobank (UKB). To support both the execution and visualization of these analyses, we also developed a web-based platform, UK-IGV (UK Biobank–Based Interpretation of Genetic Variants). Using WES data and clinical records from approximately 200,000 participants, we conducted Phenome-Wide Association Studies (PheWAS) on 65 clinical laboratory measurements. The platform enables users to explore association results not only at the population level but also across various subgroups, including age ranges, sex, and disease cohorts. Statistically significant PheWAS results are visualized through violin plots and Manhattan plots to enhance interpretability.
In addition, UK-IGV integrates external annotation sources to support functional interpretation of variants. These include functional prediction scores (SIFT, PolyPhen-2, CADD, PrimateAI, AlphaMissense), disease relevance annotations (ClinVar, OMIM), and population-specific allele frequencies (1KGP, gnomAD, UKB).
Based on the PheWAS results, we constructed feature vectors for each variant and defined phenotype- and cluster-level influence scores. We then performed unsupervised clustering using machine learning techniques to identify functionally coherent variant groups. Dominant biomarkers were determined for each cluster, allowing for characterization of functional patterns among variants. Furthermore, enrichment analysis at the dominant biomarker cluster level was implemented to provide insights into the underlying biological mechanisms of each cluster.
UK-IGV serves as an integrated platform that unifies variant–phenotype association analysis, visualization, external annotation, and clustering-based interpretation. This tool is expected to facilitate the discovery and validation of candidate genetic variants and genes in biobank-scale genomic studies.
목차 (Table of Contents)