The comparative study of multi-omics data across species is essential for understanding the diversity and commonality of genome regulation, evolution, and disease biology. To contribute to this field, this dissertation presents two major advances: (1)...
The comparative study of multi-omics data across species is essential for understanding the diversity and commonality of genome regulation, evolution, and disease biology. To contribute to this field, this dissertation presents two major advances: (1) a comprehensive functional annotation of the dog genome through multi-omics profiling, and (2) the development of a unified computational tool for the efficient retrieval and integration of multi-omics data from leading public databases.
In Chapter I, we addressed the current lack of comprehensive annotations of functional genomic elements in dogs, despite the availability of high-quality draft references from large-scale genome projects. To this end, we performed integrative next-generation sequencing of transcriptomes paired with five histone marks and DNA methylome profiling across 11 tissue types, deciphering the dog’s epigenetic code by defining distinct chromatin states, super-enhancer, and DNA methylome landscapes, and thus showed that these regions are associated with a wide range of biological functions and cell/tissue identity. In addition, we confirmed that the phenotype-associated variants are enriched in tissue-specific regulatory regions and, therefore, the tissue of origin of the variants can be traced. Ultimately, we delineated conserved and dynamic epigenomic changes at the tissue- and species-specific resolutions. Our study provides an epigenomic blueprint of the dog that can be used for comparative biology and medical research.
In Chapter II, to address challenges in accessing and integrating the rapidly expanding genomic resources, I developed Gencube, a Python-based command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types—genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources—from various leading databases such as GenBank, RefSeq, Ensembl Beta, GenArk, and Zoonomia. By enabling consistent and efficient access to diverse datasets, Gencube facilitates large-scale comparative genomics and multi-omics analyses, accelerating research efforts across species.
Collectively, this dissertation contributes both a new reference framework for comparative epigenomics in dogs and a practical informatics solution for navigating the growing landscape of multi-omics resources. These advancements aim to promote more integrative and scalable comparative biology studies and offer valuable platforms for future research into genome function, evolution, and disease.