In microbiome research, an approach based on changes in the abundance of specific microbes is widely used to analyze disease associations. However, real-world microbiome data has their unique characteristics such as sparsity, compositionality, high-di...
In microbiome research, an approach based on changes in the abundance of specific microbes is widely used to analyze disease associations. However, real-world microbiome data has their unique characteristics such as sparsity, compositionality, high-dimensionality and taxonomic hierarchy. Unfortunately, these characteristics are often not fully accounted for by conventional methods (e.g., edgeR, metagenomeSeq, ANCOM, etc.). In addition, some conventional methods may overlook correlated or biologically important taxa. As a result, such an approach is limited by their statistical models and lack biological interpretability.
To address these limitations, this study proposes a hierarchical structured component model incorporating taxonomic hierarchy for microbiome data (HisCoM-microb). The model resolves analytical challenges in microbiome data by: (1) addressing the sparsity problem through a permutation-based approach; (2) correcting the compositional constraint using centered log-ratio (CLR) transformation; and (3) incorporating taxonomic hierarchy into the model structure to reflect biological information.
We evaluated HisCoM-microb using both simulation study and real colorectal cancer (CRC) microbiome data (120 CRC patients and 172 controls, 335 OTUs). To comprehensively assess HisCoM-microb, we conducted both single-marker and multi-marker analyses. The single-marker analysis evaluated each taxon independently to assess differential abundance (e.g., edgeR, ANCOM), while the multi-marker analysis jointly analyzed multiple taxa and performed feature selection (e.g., LASSO, gLASSO). Validation of identified taxa was performed through literature review, ontology of host–microbiome interactions (OHMI) database, and taxon set enrichment analysis (TSEA) to confirm their biological relevance.
In simulation study, we constructed a simulated microbial environment to evaluate the performance of HisCoM-microb under different levels of sparsity and effect size. As the effect size increased, the advantage of HisCoM-microb became more evident, particularly under low- and high-sparsity conditions. In CRC data analysis, HisCoM-microb showed the higher true discovery rate (TDR) among the compared methods across taxonomic levels, demonstrating its ability to identify markers truly associated with the disease. Moreover, through TSEA, we further confirmed the biological relevance of identified taxa. The model successfully identified Fusobacteria linage, a representative CRC-associated lineage commonly identified by other methods, and additionally identified Coprobacillus and Coprococcus, CRC-related taxa that was not identified by the others.
HisCoM-microb implements a two-layer modeling approach that considers both multiple OTUs and their corresponding taxa to reflect biological information into the model structure. Its practical utility and biological relevance have been demonstrated with CRC data. These results highlight the potential of HisCoM-microb for biomarker discovery and suggest its applicability to a wide range of disease studies.