This dissertation introduces two novel computational tools: one for large-scale analysis of protein complex structures and the other for rapid and efficient pathogen detection. The first tool, Foldseek-Multimer, is a fast and sensitive structural alig...
This dissertation introduces two novel computational tools: one for large-scale analysis of protein complex structures and the other for rapid and efficient pathogen detection. The first tool, Foldseek-Multimer, is a fast and sensitive structural alignment framework designed to identify similar protein complexes across large-scale structural databases. It supports systematic exploration of structural diversity and offers valuable insights into protein complex function and evolutionary relationships. Recent advances in protein structure prediction, such as RoseTTAFold, AlphaFold2, and AlphaFold-Multimer, have extended modeling capabilities from individual polypeptides to protein complexes, enabling proteome-wide prediction of quaternary structures. Despite this progress, converting the vast volume of predicted structural data into meaningful biological knowledge remains a significant challenge. This demands structural comparison tools that are not only accurate but also scalable beyond the limitations of traditional methods. Foldseek-Multimer was developed to address this need. The second tool, Probeit, responds to the urgent demand for rapid pathogen detection, a need that was strongly emphasized during the COVID-19 pandemic. Probeit is a high-speed, reliable probe design tool that facilitates the timely and accurate identification of pathogens. By enabling efficient probe generation, it provides a practical solution for public health applications requiring fast diagnostics and surveillance.
This thesis is organized into five chapters. Chapter 1 provides a review of related literature. Chapter 2 introduces Foldseek-Multimer, a fast and sensitive structural aligner for protein complexes. Chapter 3 presents the detection of structurally similar protein complexes related to a CRISPR–Cas type IV-A system from a large-scale database. Chapter 4 describes Probeit, a rapid and efficient tool for probe design. Chapter 5 presents conclusions.
Chapter 2 introduces Foldseek-Multimer, a tool designed for fast and accurate alignment of protein complexes. Recent breakthroughs in computational structure prediction—such as AlphaFold-Multimer, RoseTTAFold2, and AlphaFold3—have made it possible to predict the quaternary structures of protein complexes. These tools have enabled researchers to scale up structural analysis to the level of proteome-wide complex structure prediction, generating millions of predicted assemblies. Analyzing and comparing these quaternary structures is essential for understanding structural diversity, which is largely shaped by interactions between protein chains. However, turning this DB-scaled data into meaningful biological insights requires efficient methods for structural alignment and comparison. This task remains computationally demanding with current state-of-the-art tools. In particular, determining the optimal chain matching between two protein complexes requires a factorial number of comparisons, making naive approaches infeasible at scale. To address this challenge, I developed Foldseek-Multimer, a high-speed structural aligner for protein complexes. Foldseek-Multimer is 3 to 4 orders of magnitude faster than the gold-standard method US-align, while maintaining comparable alignment quality. This speed enables the alignment of billions of complex pairs within just 11 hours. Foldseek-Multimer achieves this performance through three key innovations: (1) leveraging Foldseek for rapid chain-to-chain structural comparisons, (2) using clustered databases to reduce search complexity, and (3) representing chain-to-chain alignments as superposition vectors, which are then clustered to identify optimal chain matchings. With Foldseek-Multimer, researchers can efficiently detect structural homology across large-scale structure databases, uncover evolutionary relationships, and make functional predictions for protein complexes in previously unexplored organisms.
Chapter 3 demonstrates the capability of Foldseek-Multimer to detect structurally homologous protein complexes from large-scale databases. Identifying structural homology of a novel protein complex is essential for inferring its function. To validate Foldseek-Multimer's effectiveness, we performed the environmental CRISPR–Cas benchmark. Foldseek-Multimer completed the task in just 30 seconds, while US-align required 13 days, yet both methods identified the same five structurally similar complexes. A benchmark assessing prediction quality revealed that lower structural accuracy in the input leads to reduced TM-scores in the resulting alignments. Additionally, the accuracy of structural homology detection was shown to depend on the quality of the predicted structure of the query complex. Together, these results highlight Foldseek-Multimer's utility for rapidly and reliably identifying functional and structural homologs from large databases—an essential step in characterizing unannotated protein complexes.
Chapter 4 introduces Probeit, a rapid and efficient probe designer. During the COVID-19 pandemic, the importance of pathogen detection became prominent. To enable rapid and efficient detection, there is a growing need for a rapid and effective probe designer. I developed Probeit to meet this need and offers several key features: (1) efficient handling of large-scale input sequences through redundancy reduction, (2) design of selective probe sets that avoid host-derived sequences, (3) incorporation of thermodynamic filtering using Primer3 to ensure probe stability, and (4) generation of compact probe sets that do not fully tile all input sequences. To compare, I designed probe sets using Probeit and CATCH, a widely used probe designer, to cover all Alphainfluenzavirus CDS with avian hosts in GenBank while avoiding chicken (Gallus gallus) sequences. This benchmark highlights the efficiency of Probeit’s redundancy reduction and the compactness of the resulting probe sets.
Foldseek-Multimer and Probeit were designed to empower biological researchers in addressing real-world challenges, enabling large-scale sequence and structure alignment, and facilitating the targeted detection of biological features. I hope that these tools will accelerate new biological discoveries and make meaningful contributions to the broader scientific community.