RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    시퀀싱 리드 깊이 분포를 활용한 복제 수 변이 탐색 연구 = A Study on the CNV Detection Using Sequencing Read Depth Distribution

    한글로보기

    https://www.riss.kr/link?id=T17380514

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    차세대 시퀀싱(Next-Generation Sequencing, NGS) 기술의 비약적인 발 전은 데이터 생성 속도의 가속화와 분석 비용의 획기적인 절감을 가져왔으 며, 이는 대규모 전장 유전체 시퀀싱(Whole Genome Sequencing, WGS)을 통한 정밀한 돌연변이 검출을 가능하게 하였다. 다양한 유전체 변이 중에서 도 복제 수 변이(Copy Number Variation, CNV)는 유전자 발현 조절 및 표 현형 결정에 영향을 미치며, 수많은 유전 질환 및 암 발병 기전과 밀접한 연관이 있음이 밝혀졌다. 따라서 CNV와 관련된 연구에서는 CNV의 경계를 정확하게 정의하는 것이 필수적이다. 그러나 기존의 CNV 탐색 도구 (Calling tool)는 미리 정의된 해상도에 의존하여 분석을 수행하기 때문에, 실제 변이 구간과 탐지된 구간 사이에 불일치가 발생하거나 부정확한 경계 추정을 초래하는 한계가 있다. 이 문제를 해결하기 위해, 본 연구에서는 기존 CNV 탐색 도구로부터 얻 은 결과를 리드 정렬 데이터를 통해 개선하여 CNV 경계 정확도를 향상시 키는 개선 방법론을 제시한다. 제안된 방법론의 유효성을 검증하기 위해 1000 Genome Project의 전체 유전체 시퀀싱 데이터와 유전체 변이 데이터 베이스(DGV)의 검증된 CNV를 사용하였다. 실험 결과, 본 연구의 방법론은 기존 CNV 중 최대 약 75%까지 성공적으로 개선하고 전체 CNV의 평균 F1-score를 0.31에서 0.33으로 향상시켰다. 나아가 기존의 CNV가 탐색 도구와 길이에 상관없이 일관된 성능을 보였다. 이러한 결과는 본 연구에서 제안하는 방법론이 기존 분석 파이프라인의 한계를 보완하고 CNV 경계 정밀도를 효과적으로 향상시킴으로써, 향후 CNV 연관 질병의 병인 규명 및 대규모 집단 유전학 연구의 데이터 신뢰도를 높이는 데 기여할 수 있음을 시사한다. 본 연구의 방법론 소스 코드 및 튜토리얼 데이터는 다음 주소에 서 공개되어 이용할 수 있다: https://github.com/jkimlab/Dr.CNV.
    번역하기

    차세대 시퀀싱(Next-Generation Sequencing, NGS) 기술의 비약적인 발 전은 데이터 생성 속도의 가속화와 분석 비용의 획기적인 절감을 가져왔으 며, 이는 대규모 전장 유전체 시퀀싱(Whole Genome Sequencin...

    차세대 시퀀싱(Next-Generation Sequencing, NGS) 기술의 비약적인 발 전은 데이터 생성 속도의 가속화와 분석 비용의 획기적인 절감을 가져왔으 며, 이는 대규모 전장 유전체 시퀀싱(Whole Genome Sequencing, WGS)을 통한 정밀한 돌연변이 검출을 가능하게 하였다. 다양한 유전체 변이 중에서 도 복제 수 변이(Copy Number Variation, CNV)는 유전자 발현 조절 및 표 현형 결정에 영향을 미치며, 수많은 유전 질환 및 암 발병 기전과 밀접한 연관이 있음이 밝혀졌다. 따라서 CNV와 관련된 연구에서는 CNV의 경계를 정확하게 정의하는 것이 필수적이다. 그러나 기존의 CNV 탐색 도구 (Calling tool)는 미리 정의된 해상도에 의존하여 분석을 수행하기 때문에, 실제 변이 구간과 탐지된 구간 사이에 불일치가 발생하거나 부정확한 경계 추정을 초래하는 한계가 있다. 이 문제를 해결하기 위해, 본 연구에서는 기존 CNV 탐색 도구로부터 얻 은 결과를 리드 정렬 데이터를 통해 개선하여 CNV 경계 정확도를 향상시 키는 개선 방법론을 제시한다. 제안된 방법론의 유효성을 검증하기 위해 1000 Genome Project의 전체 유전체 시퀀싱 데이터와 유전체 변이 데이터 베이스(DGV)의 검증된 CNV를 사용하였다. 실험 결과, 본 연구의 방법론은 기존 CNV 중 최대 약 75%까지 성공적으로 개선하고 전체 CNV의 평균 F1-score를 0.31에서 0.33으로 향상시켰다. 나아가 기존의 CNV가 탐색 도구와 길이에 상관없이 일관된 성능을 보였다. 이러한 결과는 본 연구에서 제안하는 방법론이 기존 분석 파이프라인의 한계를 보완하고 CNV 경계 정밀도를 효과적으로 향상시킴으로써, 향후 CNV 연관 질병의 병인 규명 및 대규모 집단 유전학 연구의 데이터 신뢰도를 높이는 데 기여할 수 있음을 시사한다. 본 연구의 방법론 소스 코드 및 튜토리얼 데이터는 다음 주소에 서 공개되어 이용할 수 있다: https://github.com/jkimlab/Dr.CNV.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The rapid advancement of next-generation sequencing (NGS) technologies has accelerated data generation and significantly reduced analysis costs, enabling precise detection of genomic variants through large-scale whole-genome sequencing (WGS). Among various genomic variants, copy number variation (CNV) influences gene expression regulation and phenotype determination and has been shown to be closely associated with numerous genetic disorders and cancer development. Accurate detection of CNV boundaries is therefore essential in CNV- related studies. However, existing CNV calling tools rely on predefined resolution parameters, which often result in discrepancies between the true CNV boundaries and the detected regions. To address these limitations, we propose an improved approach that refines CNV boundaries by leveraging read alignment data to enhance results obtained from existing CNV callers. We validated the proposed methodology using whole-genome sequencing data from the 1000 Genomes Project and curated CNVs from the Database of Genomic Variants (DGV). The result of this study demonstrate that our method successfully refines up to approximately 75% of existing CNVs and improves the overall average F1-score from 0.31 to 0.33. Also, the refined CNVs exhibit consistent performance regardless of the calling tool or CNV length. These results indicate that this approach effectively overcomes key limitations of current CNV analysis pipelines and offers improved boundary precision. Furthermore, contributing to more reliable insights in studies of CNV-associated disease mechanisms and large- scale population genetics. The source code and tutorial data for the proposed method are publicly available at: https://github.com/jkimlab/Dr.CNV
    번역하기

    The rapid advancement of next-generation sequencing (NGS) technologies has accelerated data generation and significantly reduced analysis costs, enabling precise detection of genomic variants through large-scale whole-genome sequencing (WGS). Among va...

    The rapid advancement of next-generation sequencing (NGS) technologies has accelerated data generation and significantly reduced analysis costs, enabling precise detection of genomic variants through large-scale whole-genome sequencing (WGS). Among various genomic variants, copy number variation (CNV) influences gene expression regulation and phenotype determination and has been shown to be closely associated with numerous genetic disorders and cancer development. Accurate detection of CNV boundaries is therefore essential in CNV- related studies. However, existing CNV calling tools rely on predefined resolution parameters, which often result in discrepancies between the true CNV boundaries and the detected regions. To address these limitations, we propose an improved approach that refines CNV boundaries by leveraging read alignment data to enhance results obtained from existing CNV callers. We validated the proposed methodology using whole-genome sequencing data from the 1000 Genomes Project and curated CNVs from the Database of Genomic Variants (DGV). The result of this study demonstrate that our method successfully refines up to approximately 75% of existing CNVs and improves the overall average F1-score from 0.31 to 0.33. Also, the refined CNVs exhibit consistent performance regardless of the calling tool or CNV length. These results indicate that this approach effectively overcomes key limitations of current CNV analysis pipelines and offers improved boundary precision. Furthermore, contributing to more reliable insights in studies of CNV-associated disease mechanisms and large- scale population genetics. The source code and tutorial data for the proposed method are publicly available at: https://github.com/jkimlab/Dr.CNV

    더보기

    목차 (Table of Contents)

    • 표목차 ⅱ
    • 그림목차 ⅲ
    • 국문초록 ⅳ
    • 제1장 서론 1
    • 표목차 ⅱ
    • 그림목차 ⅲ
    • 국문초록 ⅳ
    • 제1장 서론 1
    • 제2장 개선 방법론 개발 및 연구 4
    • 제1절 CNV 경계 개선 방법론 개발 4
    • 1. 방법론 설계 및 입출력 4
    • 2. CNV 경계 개선 알고리즘 구현 6
    • 3. 알고리즘 구현 환경 8
    • 제2절 방법론 실험 및 평가 9
    • 1. 데이터세트 구성 과정 9
    • 2. 방법론 평가 설계 및 과정 11
    • 제3장 방법론 성능 평가 13
    • 제1절 정량적 개선 성능 평가 13
    • 제2절 길이 분포에 따른 성능 분석 17
    • 제4장 결론 19
    • 참고문헌 20
    • ABSTRACT 23
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼