RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Computational Methods for Protein Complex and Pathogen Detection = 단백질 복합체 및 병원체 탐지를 위한 전산학적 방법

    한글로보기

    https://www.riss.kr/link?id=T17314526

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This dissertation introduces two novel computational tools: one for large-scale analysis of protein complex structures and the other for rapid and efficient pathogen detection. The first tool, Foldseek-Multimer, is a fast and sensitive structural alignment framework designed to identify similar protein complexes across large-scale structural databases. It supports systematic exploration of structural diversity and offers valuable insights into protein complex function and evolutionary relationships. Recent advances in protein structure prediction, such as RoseTTAFold, AlphaFold2, and AlphaFold-Multimer, have extended modeling capabilities from individual polypeptides to protein complexes, enabling proteome-wide prediction of quaternary structures. Despite this progress, converting the vast volume of predicted structural data into meaningful biological knowledge remains a significant challenge. This demands structural comparison tools that are not only accurate but also scalable beyond the limitations of traditional methods. Foldseek-Multimer was developed to address this need. The second tool, Probeit, responds to the urgent demand for rapid pathogen detection, a need that was strongly emphasized during the COVID-19 pandemic. Probeit is a high-speed, reliable probe design tool that facilitates the timely and accurate identification of pathogens. By enabling efficient probe generation, it provides a practical solution for public health applications requiring fast diagnostics and surveillance.

    This thesis is organized into five chapters. Chapter 1 provides a review of related literature. Chapter 2 introduces Foldseek-Multimer, a fast and sensitive structural aligner for protein complexes. Chapter 3 presents the detection of structurally similar protein complexes related to a CRISPR–Cas type IV-A system from a large-scale database. Chapter 4 describes Probeit, a rapid and efficient tool for probe design. Chapter 5 presents conclusions.

    Chapter 2 introduces Foldseek-Multimer, a tool designed for fast and accurate alignment of protein complexes. Recent breakthroughs in computational structure prediction—such as AlphaFold-Multimer, RoseTTAFold2, and AlphaFold3—have made it possible to predict the quaternary structures of protein complexes. These tools have enabled researchers to scale up structural analysis to the level of proteome-wide complex structure prediction, generating millions of predicted assemblies. Analyzing and comparing these quaternary structures is essential for understanding structural diversity, which is largely shaped by interactions between protein chains. However, turning this DB-scaled data into meaningful biological insights requires efficient methods for structural alignment and comparison. This task remains computationally demanding with current state-of-the-art tools. In particular, determining the optimal chain matching between two protein complexes requires a factorial number of comparisons, making naive approaches infeasible at scale. To address this challenge, I developed Foldseek-Multimer, a high-speed structural aligner for protein complexes. Foldseek-Multimer is 3 to 4 orders of magnitude faster than the gold-standard method US-align, while maintaining comparable alignment quality. This speed enables the alignment of billions of complex pairs within just 11 hours. Foldseek-Multimer achieves this performance through three key innovations: (1) leveraging Foldseek for rapid chain-to-chain structural comparisons, (2) using clustered databases to reduce search complexity, and (3) representing chain-to-chain alignments as superposition vectors, which are then clustered to identify optimal chain matchings. With Foldseek-Multimer, researchers can efficiently detect structural homology across large-scale structure databases, uncover evolutionary relationships, and make functional predictions for protein complexes in previously unexplored organisms.

    Chapter 3 demonstrates the capability of Foldseek-Multimer to detect structurally homologous protein complexes from large-scale databases. Identifying structural homology of a novel protein complex is essential for inferring its function. To validate Foldseek-Multimer's effectiveness, we performed the environmental CRISPR–Cas benchmark. Foldseek-Multimer completed the task in just 30 seconds, while US-align required 13 days, yet both methods identified the same five structurally similar complexes. A benchmark assessing prediction quality revealed that lower structural accuracy in the input leads to reduced TM-scores in the resulting alignments. Additionally, the accuracy of structural homology detection was shown to depend on the quality of the predicted structure of the query complex. Together, these results highlight Foldseek-Multimer's utility for rapidly and reliably identifying functional and structural homologs from large databases—an essential step in characterizing unannotated protein complexes.

    Chapter 4 introduces Probeit, a rapid and efficient probe designer. During the COVID-19 pandemic, the importance of pathogen detection became prominent. To enable rapid and efficient detection, there is a growing need for a rapid and effective probe designer. I developed Probeit to meet this need and offers several key features: (1) efficient handling of large-scale input sequences through redundancy reduction, (2) design of selective probe sets that avoid host-derived sequences, (3) incorporation of thermodynamic filtering using Primer3 to ensure probe stability, and (4) generation of compact probe sets that do not fully tile all input sequences. To compare, I designed probe sets using Probeit and CATCH, a widely used probe designer, to cover all Alphainfluenzavirus CDS with avian hosts in GenBank while avoiding chicken (Gallus gallus) sequences. This benchmark highlights the efficiency of Probeit’s redundancy reduction and the compactness of the resulting probe sets.

    Foldseek-Multimer and Probeit were designed to empower biological researchers in addressing real-world challenges, enabling large-scale sequence and structure alignment, and facilitating the targeted detection of biological features. I hope that these tools will accelerate new biological discoveries and make meaningful contributions to the broader scientific community.
    번역하기

    This dissertation introduces two novel computational tools: one for large-scale analysis of protein complex structures and the other for rapid and efficient pathogen detection. The first tool, Foldseek-Multimer, is a fast and sensitive structural alig...

    This dissertation introduces two novel computational tools: one for large-scale analysis of protein complex structures and the other for rapid and efficient pathogen detection. The first tool, Foldseek-Multimer, is a fast and sensitive structural alignment framework designed to identify similar protein complexes across large-scale structural databases. It supports systematic exploration of structural diversity and offers valuable insights into protein complex function and evolutionary relationships. Recent advances in protein structure prediction, such as RoseTTAFold, AlphaFold2, and AlphaFold-Multimer, have extended modeling capabilities from individual polypeptides to protein complexes, enabling proteome-wide prediction of quaternary structures. Despite this progress, converting the vast volume of predicted structural data into meaningful biological knowledge remains a significant challenge. This demands structural comparison tools that are not only accurate but also scalable beyond the limitations of traditional methods. Foldseek-Multimer was developed to address this need. The second tool, Probeit, responds to the urgent demand for rapid pathogen detection, a need that was strongly emphasized during the COVID-19 pandemic. Probeit is a high-speed, reliable probe design tool that facilitates the timely and accurate identification of pathogens. By enabling efficient probe generation, it provides a practical solution for public health applications requiring fast diagnostics and surveillance.

    This thesis is organized into five chapters. Chapter 1 provides a review of related literature. Chapter 2 introduces Foldseek-Multimer, a fast and sensitive structural aligner for protein complexes. Chapter 3 presents the detection of structurally similar protein complexes related to a CRISPR–Cas type IV-A system from a large-scale database. Chapter 4 describes Probeit, a rapid and efficient tool for probe design. Chapter 5 presents conclusions.

    Chapter 2 introduces Foldseek-Multimer, a tool designed for fast and accurate alignment of protein complexes. Recent breakthroughs in computational structure prediction—such as AlphaFold-Multimer, RoseTTAFold2, and AlphaFold3—have made it possible to predict the quaternary structures of protein complexes. These tools have enabled researchers to scale up structural analysis to the level of proteome-wide complex structure prediction, generating millions of predicted assemblies. Analyzing and comparing these quaternary structures is essential for understanding structural diversity, which is largely shaped by interactions between protein chains. However, turning this DB-scaled data into meaningful biological insights requires efficient methods for structural alignment and comparison. This task remains computationally demanding with current state-of-the-art tools. In particular, determining the optimal chain matching between two protein complexes requires a factorial number of comparisons, making naive approaches infeasible at scale. To address this challenge, I developed Foldseek-Multimer, a high-speed structural aligner for protein complexes. Foldseek-Multimer is 3 to 4 orders of magnitude faster than the gold-standard method US-align, while maintaining comparable alignment quality. This speed enables the alignment of billions of complex pairs within just 11 hours. Foldseek-Multimer achieves this performance through three key innovations: (1) leveraging Foldseek for rapid chain-to-chain structural comparisons, (2) using clustered databases to reduce search complexity, and (3) representing chain-to-chain alignments as superposition vectors, which are then clustered to identify optimal chain matchings. With Foldseek-Multimer, researchers can efficiently detect structural homology across large-scale structure databases, uncover evolutionary relationships, and make functional predictions for protein complexes in previously unexplored organisms.

    Chapter 3 demonstrates the capability of Foldseek-Multimer to detect structurally homologous protein complexes from large-scale databases. Identifying structural homology of a novel protein complex is essential for inferring its function. To validate Foldseek-Multimer's effectiveness, we performed the environmental CRISPR–Cas benchmark. Foldseek-Multimer completed the task in just 30 seconds, while US-align required 13 days, yet both methods identified the same five structurally similar complexes. A benchmark assessing prediction quality revealed that lower structural accuracy in the input leads to reduced TM-scores in the resulting alignments. Additionally, the accuracy of structural homology detection was shown to depend on the quality of the predicted structure of the query complex. Together, these results highlight Foldseek-Multimer's utility for rapidly and reliably identifying functional and structural homologs from large databases—an essential step in characterizing unannotated protein complexes.

    Chapter 4 introduces Probeit, a rapid and efficient probe designer. During the COVID-19 pandemic, the importance of pathogen detection became prominent. To enable rapid and efficient detection, there is a growing need for a rapid and effective probe designer. I developed Probeit to meet this need and offers several key features: (1) efficient handling of large-scale input sequences through redundancy reduction, (2) design of selective probe sets that avoid host-derived sequences, (3) incorporation of thermodynamic filtering using Primer3 to ensure probe stability, and (4) generation of compact probe sets that do not fully tile all input sequences. To compare, I designed probe sets using Probeit and CATCH, a widely used probe designer, to cover all Alphainfluenzavirus CDS with avian hosts in GenBank while avoiding chicken (Gallus gallus) sequences. This benchmark highlights the efficiency of Probeit’s redundancy reduction and the compactness of the resulting probe sets.

    Foldseek-Multimer and Probeit were designed to empower biological researchers in addressing real-world challenges, enabling large-scale sequence and structure alignment, and facilitating the targeted detection of biological features. I hope that these tools will accelerate new biological discoveries and make meaningful contributions to the broader scientific community.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    이 논문은 두 가지 새로운 계산 도구를 소개한다: 하나는 단백질 복합체 구조의 대규모 분석을 위한 것이고, 다른 하나는 신속하고 효율적인 병원체 검출을 위한 것이다. 첫 번째 도구인 Foldseek-Multimer는 대규모 구조 데이터베이스에서 유사한 단백질 복합체를 식별하도록 설계된 빠르고 민감한 구조 정렬 프레임워크이다. 이는 구조적 다양성에 관한 체계적인 탐구를 지원하며, 단백질 복합체 기능과 진화적 관계에 대한 귀중한 통찰력을 제공한다. 최근 단백질 구조 예측의 발전, 예를 들어 RoseTTAFold, AlphaFold2, AlphaFold-Multimer는 개별 폴리펩타이드에서 단백질 복합체로 모델링 기능을 확장하여 4차 구조의 프로테옴 전반을 예측할 수 있게 하였다. 이러한 발전에도 불구하고, 방대한 양의 예측된 구조 데이터를 의미 있는 생물학적 지식으로 변환하는 것은 여전히 중요한 도전 과제로 남아 있다. 이를 위해서는 정확할 뿐만 아니라 전통적인 방법의 한계를 넘어서는 확장 가능한 구조 비교 도구가 요구된다. 이러한 요구를 해결하기 위해 우리는 Foldseek-Multimer를 개발했다. 두 번째 도구인 Probeit은 COVID-19 팬데믹 동안 강력하게 강조되었던 신속한 병원균 검출에 대한 긴급한 수요에 대응한다. Probeit은 병원균을 적시에 정확하게 식별할 수 있도록 지원하는 빠르고 신뢰할 수 있는 프로브 설계 도구로, 효율적인 프로브 생성을 가능하게 함으로써 빠른 진단과 감시가 필요한 공중 보건 애플리케이션에 실질적인 해결책을 제공한다.

    본 논문은 총 5장으로 구성되었다. 제1장은 관련 문헌들의 검토를 제공한다. 제2장은 빠르고 민감한 단백질 복합체 구조 정렬기인 Foldseek-Multimer를 소개한다. 제3장은 Foldseek-Multimer를 이용해 대규모 데이터베이스로부터 CRISPR–Cas type IV-A 시스템과 구조적으로 유사한 단백질 복합체를 탐지한 사례를 제시한다. 제4장은 빠르고 효율적인 프로브 디자이너인 Probeit을 소개한다. 제5장은 본 연구의 결론을 제시한다.

    제2장은 Foldseek-Multimer를 소개한다. AlphaFold-Multimer, RoseTTAFold2, AlphaFold3와 같은 최신 계산 구조 예측 기술의 발전은 단백질 복합체의 4차 구조(quaternary structure) 예측을 가능하게 만들었다. 이러한 도구들은 연구자들이 수백만 개에 달하는 예측 구조를 생성할 수 있도록 하여, 단백질체(proteome) 전체 수준에서의 복합체 구조 분석을 가능하게 한다. 이러한 4차 구조들을 분석하고 비교하는 것은 단백질 사슬 간의 상호작용에 의해 형성되는 구조적 다양성을 이해하는 데 필수적이다. 하지만, 이 데이터베이스 규모의 구조 예측 데이터를 생물학적으로 의미 있는 발견으로 전환하려면, 효율적인 구조 정렬 및 비교 방법이 필요하다. 특히 두 단백질 복합체 간의 최적의 체인 매칭을 결정하는 데는 팩토리얼 수준의 계산량이 요구되기 때문에, 기존의 접근 방식으로는 대규모 분석이 현실적으로 어렵다. 이러한 문제를 해결하기 위해, 나는 Foldseek-Multimer라는 단백질 복합체 정렬 도구를 개발했다. Foldseek-Multimer는 골드-스탠다드 도구인 US-align보다 3-4자릿수 빠르며, 정렬 품질은 유사한 수준을 유지한다. 이를 통해 수십억 쌍의 복합체를 단 11시간 안에 처리할 수 있다. Foldseek-Multimer는 다음의 세 가지 핵심 전략을 통해 성능을 극적으로 향상시켰다: (1) 빠른 체인 간 구조 비교를 위해 Foldseek을 활용하고, (2) 클러스터링된 데이터베이스를 사용하여 탐색 효율을 최적화하며, (3) 체인 간 정렬을 중첩 벡터(superposition vector)로 표현한 후, 이를 클러스터링하여 체인 매칭을 식별한다. Foldseek-Multimer를 통해 연구자들은 대규모 구조 데이터베이스에서 구조적 상동성을 빠르게 탐지하고, 진화적 관계를 밝혀내며, 기존에 탐색되지 않았던 생물종의 단백질 복합체에 대해 기능 예측을 수행할 수 있게 될 것이다.

    제3장은 Foldseek-Multimer를 사용하여 대규모 데이터베이스로부터 구조적으로 유사한 단백질 복합체를 탐지하는 능력을 보여준다. 새로운 단백질 복합체의 기능을 예측하기 위해서는 구조적으로 유사한 복합체를 탐지하는 것이 중요하다. 이를 검증하기 위해, Environmental CRISPR–Cas 벤치마크를 수행하였다. Foldseek-Multimer는 30초 만에 탐지를 완료했한 반면, US-align은 같은 작업을 수행하는 데 13일이 소요되었다. 두 도구 모두 동일한 5개의 유사 복합체를 탐지하였다. 또한, 예측된 구조의 품질이 낮을수록 TM-score가 하락한다는 벤치마크 결과를 통해, 구조 예측 품질이 구조 정렬의 정확도에 미치는 영향을 확인하였다. 이 결과는 Foldseek-Multimer가 알려지지 않은 복합체의 기능적, 구조적 유사체를 대규모 데이터베이스에서 신속하고 신뢰성 있게 탐지할 수 있는 도구임을 보여준다.

    제4장은 빠르고 효율적인 프로브 설계 도구인 Probeit을 소개한다. COVID-19 팬데믹 동안 병원체 탐지의 중요성이 강조되었고, 이를 위한 빠르고 효과적인 프로브 디자이너의 필요성이 대두되었다. 이를 위해 개발된 Probeit은 다음과 같은 특징을 가진다: (1) MMseqs2-linclust를 통한 대규모 입력 데이터의 효율적 처리(Redundancy reduction), (2) host 유래 서열을 회피하는 선택적 probe set의 설계, (3) Primer3를 활용한 열역학적 안정성 필터링을 통한 프로브 선택, (4) 모든 입력 서열을 완전히 커버하지 않으면서도 실험적으로 유의미한 compact probe set을 생성. 본 연구에서는 GenBank의 조류 인플루엔자 A형 바이러스의 CDS를 대상으로, 닭(Gallus gallus) 유래 서열을 회피하도록 하여 Probeit과 널리 사용되는 CATCH 도구를 이용해 probe set을 설계하였다. 벤치마크 결과는 Probeit이 redundancy reduction을 통해 빠르고 작고 효율적인 probe set을 생성할 수 있음을 보여준다.

    Foldseek-Multimer와 Probeit은 생물학 연구자들이 실제 문제를 해결하고, 대규모 서열 및 구조 정렬을 수행하며, 생물학적 특징을 보다 쉽게 탐지할 수 있도록 설계된 도구이다. 나는 이러한 도구들이 새로운 생물학적 발견을 가속화하고, 더 넓은 과학 커뮤니티에 이바지할 수 있기를 바란다.
    번역하기

    이 논문은 두 가지 새로운 계산 도구를 소개한다: 하나는 단백질 복합체 구조의 대규모 분석을 위한 것이고, 다른 하나는 신속하고 효율적인 병원체 검출을 위한 것이다. 첫 번째 도구인 Folds...

    이 논문은 두 가지 새로운 계산 도구를 소개한다: 하나는 단백질 복합체 구조의 대규모 분석을 위한 것이고, 다른 하나는 신속하고 효율적인 병원체 검출을 위한 것이다. 첫 번째 도구인 Foldseek-Multimer는 대규모 구조 데이터베이스에서 유사한 단백질 복합체를 식별하도록 설계된 빠르고 민감한 구조 정렬 프레임워크이다. 이는 구조적 다양성에 관한 체계적인 탐구를 지원하며, 단백질 복합체 기능과 진화적 관계에 대한 귀중한 통찰력을 제공한다. 최근 단백질 구조 예측의 발전, 예를 들어 RoseTTAFold, AlphaFold2, AlphaFold-Multimer는 개별 폴리펩타이드에서 단백질 복합체로 모델링 기능을 확장하여 4차 구조의 프로테옴 전반을 예측할 수 있게 하였다. 이러한 발전에도 불구하고, 방대한 양의 예측된 구조 데이터를 의미 있는 생물학적 지식으로 변환하는 것은 여전히 중요한 도전 과제로 남아 있다. 이를 위해서는 정확할 뿐만 아니라 전통적인 방법의 한계를 넘어서는 확장 가능한 구조 비교 도구가 요구된다. 이러한 요구를 해결하기 위해 우리는 Foldseek-Multimer를 개발했다. 두 번째 도구인 Probeit은 COVID-19 팬데믹 동안 강력하게 강조되었던 신속한 병원균 검출에 대한 긴급한 수요에 대응한다. Probeit은 병원균을 적시에 정확하게 식별할 수 있도록 지원하는 빠르고 신뢰할 수 있는 프로브 설계 도구로, 효율적인 프로브 생성을 가능하게 함으로써 빠른 진단과 감시가 필요한 공중 보건 애플리케이션에 실질적인 해결책을 제공한다.

    본 논문은 총 5장으로 구성되었다. 제1장은 관련 문헌들의 검토를 제공한다. 제2장은 빠르고 민감한 단백질 복합체 구조 정렬기인 Foldseek-Multimer를 소개한다. 제3장은 Foldseek-Multimer를 이용해 대규모 데이터베이스로부터 CRISPR–Cas type IV-A 시스템과 구조적으로 유사한 단백질 복합체를 탐지한 사례를 제시한다. 제4장은 빠르고 효율적인 프로브 디자이너인 Probeit을 소개한다. 제5장은 본 연구의 결론을 제시한다.

    제2장은 Foldseek-Multimer를 소개한다. AlphaFold-Multimer, RoseTTAFold2, AlphaFold3와 같은 최신 계산 구조 예측 기술의 발전은 단백질 복합체의 4차 구조(quaternary structure) 예측을 가능하게 만들었다. 이러한 도구들은 연구자들이 수백만 개에 달하는 예측 구조를 생성할 수 있도록 하여, 단백질체(proteome) 전체 수준에서의 복합체 구조 분석을 가능하게 한다. 이러한 4차 구조들을 분석하고 비교하는 것은 단백질 사슬 간의 상호작용에 의해 형성되는 구조적 다양성을 이해하는 데 필수적이다. 하지만, 이 데이터베이스 규모의 구조 예측 데이터를 생물학적으로 의미 있는 발견으로 전환하려면, 효율적인 구조 정렬 및 비교 방법이 필요하다. 특히 두 단백질 복합체 간의 최적의 체인 매칭을 결정하는 데는 팩토리얼 수준의 계산량이 요구되기 때문에, 기존의 접근 방식으로는 대규모 분석이 현실적으로 어렵다. 이러한 문제를 해결하기 위해, 나는 Foldseek-Multimer라는 단백질 복합체 정렬 도구를 개발했다. Foldseek-Multimer는 골드-스탠다드 도구인 US-align보다 3-4자릿수 빠르며, 정렬 품질은 유사한 수준을 유지한다. 이를 통해 수십억 쌍의 복합체를 단 11시간 안에 처리할 수 있다. Foldseek-Multimer는 다음의 세 가지 핵심 전략을 통해 성능을 극적으로 향상시켰다: (1) 빠른 체인 간 구조 비교를 위해 Foldseek을 활용하고, (2) 클러스터링된 데이터베이스를 사용하여 탐색 효율을 최적화하며, (3) 체인 간 정렬을 중첩 벡터(superposition vector)로 표현한 후, 이를 클러스터링하여 체인 매칭을 식별한다. Foldseek-Multimer를 통해 연구자들은 대규모 구조 데이터베이스에서 구조적 상동성을 빠르게 탐지하고, 진화적 관계를 밝혀내며, 기존에 탐색되지 않았던 생물종의 단백질 복합체에 대해 기능 예측을 수행할 수 있게 될 것이다.

    제3장은 Foldseek-Multimer를 사용하여 대규모 데이터베이스로부터 구조적으로 유사한 단백질 복합체를 탐지하는 능력을 보여준다. 새로운 단백질 복합체의 기능을 예측하기 위해서는 구조적으로 유사한 복합체를 탐지하는 것이 중요하다. 이를 검증하기 위해, Environmental CRISPR–Cas 벤치마크를 수행하였다. Foldseek-Multimer는 30초 만에 탐지를 완료했한 반면, US-align은 같은 작업을 수행하는 데 13일이 소요되었다. 두 도구 모두 동일한 5개의 유사 복합체를 탐지하였다. 또한, 예측된 구조의 품질이 낮을수록 TM-score가 하락한다는 벤치마크 결과를 통해, 구조 예측 품질이 구조 정렬의 정확도에 미치는 영향을 확인하였다. 이 결과는 Foldseek-Multimer가 알려지지 않은 복합체의 기능적, 구조적 유사체를 대규모 데이터베이스에서 신속하고 신뢰성 있게 탐지할 수 있는 도구임을 보여준다.

    제4장은 빠르고 효율적인 프로브 설계 도구인 Probeit을 소개한다. COVID-19 팬데믹 동안 병원체 탐지의 중요성이 강조되었고, 이를 위한 빠르고 효과적인 프로브 디자이너의 필요성이 대두되었다. 이를 위해 개발된 Probeit은 다음과 같은 특징을 가진다: (1) MMseqs2-linclust를 통한 대규모 입력 데이터의 효율적 처리(Redundancy reduction), (2) host 유래 서열을 회피하는 선택적 probe set의 설계, (3) Primer3를 활용한 열역학적 안정성 필터링을 통한 프로브 선택, (4) 모든 입력 서열을 완전히 커버하지 않으면서도 실험적으로 유의미한 compact probe set을 생성. 본 연구에서는 GenBank의 조류 인플루엔자 A형 바이러스의 CDS를 대상으로, 닭(Gallus gallus) 유래 서열을 회피하도록 하여 Probeit과 널리 사용되는 CATCH 도구를 이용해 probe set을 설계하였다. 벤치마크 결과는 Probeit이 redundancy reduction을 통해 빠르고 작고 효율적인 probe set을 생성할 수 있음을 보여준다.

    Foldseek-Multimer와 Probeit은 생물학 연구자들이 실제 문제를 해결하고, 대규모 서열 및 구조 정렬을 수행하며, 생물학적 특징을 보다 쉽게 탐지할 수 있도록 설계된 도구이다. 나는 이러한 도구들이 새로운 생물학적 발견을 가속화하고, 더 넓은 과학 커뮤니티에 이바지할 수 있기를 바란다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • List of Figures ix
    • List of Tables xii
    • Abbreviations xiv
    • Chapter 1 General Introduction 1
    • Abstract i
    • List of Figures ix
    • List of Tables xii
    • Abbreviations xiv
    • Chapter 1 General Introduction 1
    • 1.1 Recent sequence analysis tools 1
    • 1.1.1 BLAST 1
    • 1.1.2 MMseqs2 2
    • 1.1.3 MMseqs2-Linclust 2
    • 1.2 Protein structure prediction tools 2
    • 1.2.1 AlphaFold2 3
    • 1.2.2 AlphaFold-Multimer 3
    • 1.2.3 AlphaFold3 4
    • 1.2.4 RoseTTAFold 4
    • 1.2.5 RoseTTAFold2 4
    • 1.2.6 ColabFold 4
    • 1.3 Protein structural alignment tools 5
    • 1.3.1 How to evaluate protein structural alignments 5
    • 1.3.2 DALI 6
    • 1.3.3 CE 7
    • 1.3.4 TMalign 7
    • 1.3.5 Kpax 7
    • 1.3.6 Foldseek 8
    • 1.4 Protein complex structural alignment tools 11
    • 1.4.1 Difficult of protein complex structural alignment 11
    • 1.4.2 QSalign 12
    • 1.4.3 US-align, the previous gold-standard complex structural aligner 12
    • 1.5 Clustering algorithms 13
    • 1.5.1 K-means clustering 13
    • 1.5.2 K-nearest neighbors clustering 13
    • 1.5.3 Density-based spatial clustering of applications with noise 14
    • 1.5.4 Hierarchical clustering 15
    • 1.6 Probe and probe design 16
    • 1.6.1 Probe: definition and applications 16
    • 1.6.2 Efficient probe sets and the set cover problem 19
    • 1.6.3 Review of previous probe designers 19
    • 1.7 Objectives of this study 20
    • Chapter 2 Foldseek-Multimer and Protein Complex Structural Alignment 22
    • 2.1 Introduction 23
    • 2.2 Algorithm 25
    • 2.2.1 Overview 25
    • 2.2.2 Input 26
    • 2.2.3 Chain-to-chain alignments 27
    • 2.2.4 Chain-to-chain superposition vectors 27
    • 2.2.5 Chain-to-chain clustering 29
    • 2.2.6 TM-score computation 32
    • 2.2.7 Utilizing clustered databases 32
    • 2.3 Materials and methods 33
    • 2.3.1 Materials 33
    • 2.3.2 Determination of hard coded parameters 34
    • 2.3.3 Similar pairs search 37
    • 2.3.4 Database search benchmark 38
    • 2.3.5 Comparison to QSalign on 3DComplexV7 39
    • 2.4 Results 40
    • 2.4.1 Similar pairs search result 40
    • 2.4.2 Database search result 41
    • 2.4.3 Comparison to QSalign on 3DComplexV7 result 41
    • 2.5 Discussion 42
    • 2.5.1 Further studies 42
    • 2.5.2 Summary 43
    • 2.6 Availability 44
    • 2.6.1 On Foldseek-Multimer web server 44
    • 2.6.2 On BFMD resource 45
    • Chapter 3 Detection of Structurally Homologous Protein Complexes from Large-scale Database 55
    • 3.1 Introduction 56
    • 3.2 On CRISPR system 57
    • 3.3 Materials and methods 58
    • 3.3.1 Materials 58
    • 3.3.2 Environmental CRISPR-Cas benchmark 59
    • 3.3.3 Effect of complex structure prediction quality 60
    • 3.3.4 Homology detection via sequence alignment 61
    • 3.4 Results 61
    • 3.4.1 Environmental CRISPR-Cas benchmark result 61
    • 3.4.2 Effect of complex structure prediction quality 62
    • 3.4.3 Protein complex homology detection via sequence alignment 63
    • 3.5 Discussion 63
    • 3.5.1 Further studies 63
    • 3.5.2 Summary 64
    • 3.6 Availability 65
    • Chapter 4 Probeit and Pathogen Detection 71
    • 4.1 Introduction 72
    • 4.2 Algorithm 73
    • 4.2.1 Overview 73
    • 4.2.2 Positive sequences and negative sequences 74
    • 4.2.3 Redundancy reduction 74
    • 4.2.4 Extracting K-mers from redundancy reduced positive sequence 74
    • 4.2.5 Removing negative k-mers 75
    • 4.2.6 Selecting suitable k-mers 75
    • 4.2.7 Computation of input sequence mappability 77
    • 4.2.8 Construction of ligation probe sets 77
    • 4.2.9 Construction of capture probe sets 77
    • 4.3 Materials and methods 78
    • 4.3.1 Materials 78
    • 4.3.2 Benchmark 78
    • 4.3.3 Commands 79
    • 4.4 Results 79
    • 4.5 Discussion 80
    • 4.5.1 Limitations of CATCH 80
    • 4.5.2 Possible failure of Primer3 81
    • 4.5.3 Further studies 82
    • 4.5.4 Summary 83
    • 4.6 Availability 84
    • Chapter 5 Conclusion 88
    • Appendix A Supplementary Figures of Similar Pairs Benchmark 90
    • Appendix B Supplementary Figures of Environmental CRISPR-cas Benchmark 98
    • Bibliography 101
    • 국문초록 117
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼