RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Hierarchical Structured Component Model for Microbiome Data Incorporating Taxonomic Hierarchy: Application to Colorectal Cancer = 분류체계 정보를 내재한 마이크로바이옴 데이터 특화 계층적 구조 성분 모델: 대장암 적용 연구

    한글로보기

    https://www.riss.kr/link?id=T17450395

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In microbiome research, an approach based on changes in the abundance of specific microbes is widely used to analyze disease associations. However, real-world microbiome data has their unique characteristics such as sparsity, compositionality, high-dimensionality and taxonomic hierarchy. Unfortunately, these characteristics are often not fully accounted for by conventional methods (e.g., edgeR, metagenomeSeq, ANCOM, etc.). In addition, some conventional methods may overlook correlated or biologically important taxa. As a result, such an approach is limited by their statistical models and lack biological interpretability.
    To address these limitations, this study proposes a hierarchical structured component model incorporating taxonomic hierarchy for microbiome data (HisCoM-microb). The model resolves analytical challenges in microbiome data by: (1) addressing the sparsity problem through a permutation-based approach; (2) correcting the compositional constraint using centered log-ratio (CLR) transformation; and (3) incorporating taxonomic hierarchy into the model structure to reflect biological information.
    We evaluated HisCoM-microb using both simulation study and real colorectal cancer (CRC) microbiome data (120 CRC patients and 172 controls, 335 OTUs). To comprehensively assess HisCoM-microb, we conducted both single-marker and multi-marker analyses. The single-marker analysis evaluated each taxon independently to assess differential abundance (e.g., edgeR, ANCOM), while the multi-marker analysis jointly analyzed multiple taxa and performed feature selection (e.g., LASSO, gLASSO). Validation of identified taxa was performed through literature review, ontology of host–microbiome interactions (OHMI) database, and taxon set enrichment analysis (TSEA) to confirm their biological relevance.
    In simulation study, we constructed a simulated microbial environment to evaluate the performance of HisCoM-microb under different levels of sparsity and effect size. As the effect size increased, the advantage of HisCoM-microb became more evident, particularly under low- and high-sparsity conditions. In CRC data analysis, HisCoM-microb showed the higher true discovery rate (TDR) among the compared methods across taxonomic levels, demonstrating its ability to identify markers truly associated with the disease. Moreover, through TSEA, we further confirmed the biological relevance of identified taxa. The model successfully identified Fusobacteria linage, a representative CRC-associated lineage commonly identified by other methods, and additionally identified Coprobacillus and Coprococcus, CRC-related taxa that was not identified by the others.
    HisCoM-microb implements a two-layer modeling approach that considers both multiple OTUs and their corresponding taxa to reflect biological information into the model structure. Its practical utility and biological relevance have been demonstrated with CRC data. These results highlight the potential of HisCoM-microb for biomarker discovery and suggest its applicability to a wide range of disease studies.
    번역하기

    In microbiome research, an approach based on changes in the abundance of specific microbes is widely used to analyze disease associations. However, real-world microbiome data has their unique characteristics such as sparsity, compositionality, high-di...

    In microbiome research, an approach based on changes in the abundance of specific microbes is widely used to analyze disease associations. However, real-world microbiome data has their unique characteristics such as sparsity, compositionality, high-dimensionality and taxonomic hierarchy. Unfortunately, these characteristics are often not fully accounted for by conventional methods (e.g., edgeR, metagenomeSeq, ANCOM, etc.). In addition, some conventional methods may overlook correlated or biologically important taxa. As a result, such an approach is limited by their statistical models and lack biological interpretability.
    To address these limitations, this study proposes a hierarchical structured component model incorporating taxonomic hierarchy for microbiome data (HisCoM-microb). The model resolves analytical challenges in microbiome data by: (1) addressing the sparsity problem through a permutation-based approach; (2) correcting the compositional constraint using centered log-ratio (CLR) transformation; and (3) incorporating taxonomic hierarchy into the model structure to reflect biological information.
    We evaluated HisCoM-microb using both simulation study and real colorectal cancer (CRC) microbiome data (120 CRC patients and 172 controls, 335 OTUs). To comprehensively assess HisCoM-microb, we conducted both single-marker and multi-marker analyses. The single-marker analysis evaluated each taxon independently to assess differential abundance (e.g., edgeR, ANCOM), while the multi-marker analysis jointly analyzed multiple taxa and performed feature selection (e.g., LASSO, gLASSO). Validation of identified taxa was performed through literature review, ontology of host–microbiome interactions (OHMI) database, and taxon set enrichment analysis (TSEA) to confirm their biological relevance.
    In simulation study, we constructed a simulated microbial environment to evaluate the performance of HisCoM-microb under different levels of sparsity and effect size. As the effect size increased, the advantage of HisCoM-microb became more evident, particularly under low- and high-sparsity conditions. In CRC data analysis, HisCoM-microb showed the higher true discovery rate (TDR) among the compared methods across taxonomic levels, demonstrating its ability to identify markers truly associated with the disease. Moreover, through TSEA, we further confirmed the biological relevance of identified taxa. The model successfully identified Fusobacteria linage, a representative CRC-associated lineage commonly identified by other methods, and additionally identified Coprobacillus and Coprococcus, CRC-related taxa that was not identified by the others.
    HisCoM-microb implements a two-layer modeling approach that considers both multiple OTUs and their corresponding taxa to reflect biological information into the model structure. Its practical utility and biological relevance have been demonstrated with CRC data. These results highlight the potential of HisCoM-microb for biomarker discovery and suggest its applicability to a wide range of disease studies.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    마이크로바이옴 연구에서 특정 미생물의 풍부도(abundance)
    변화를 기반으로 한 접근법은 질병 연관성을 분석하는 데 널리
    사용되고 있다. 한편, 실제 마이크로바이옴 데이터는 희소성(sparsity),
    구성비 제약성(compositionality), 고차원성(high dimensionality),
    그리고 분류학적 계층 구조(taxonomic hierarchy)와 같은 고유한
    특성을 가진다. 그러나 이러한 특성들은 기존의 통계적 방법들(e.g.,
    edgeR, metagenomeSeq, ANCOM 등)에서 충분히 반영되지 못한다.
    또한 일부 기존 방법은 다른 균(taxa)과 상관성으로 인해 제외하는 등
    생물학적으로 중요한 균(taxa)의 발견을 놓칠 수 있다. 그 결과,
    이러한 접근법들은 마이크로바이옴 데이터에 적용한 기존 통계 모델의
    한계로 인해 생물학적 해석 가능성이 제한되는 문제를 갖는다.
    이러한 한계를 보완하기 위해 본 연구는 마이크로바이옴 데이터를
    위한 계층적 구조 성분 모델(hierarchical structured component
    model incorporating taxonomic hierarchy for microbiome data,
    HisCoM-microb)을 제안한다. 본 모델은 다음 세 가지 방식으로
    마이크로바이옴 데이터의 분석적 문제를 해결하도록 설계되었다: (1)
    순열 기반(permutation-based) 접근법을 통해 희소성 문제를
    완화하고, (2) 중심화 로그비(centered log-ratio, CLR) 변환을
    사용하여 구성비 제약을 보정하며, (3) 분류학적 계층 정보를 모델
    구조에 내재함으로써 생물학적 정보를 반영한다.
    본 연구는 HisCoM-microb 을 시뮬레이션 데이터와 실제
    대장암(colorectal cancer, CRC) 마이크로바이옴 데이터(120 명의
    CRC 환자와 172 명의 대조군, 총 335 OTU)를 이용해 평가하였다.
    모델의 성능을 종합적으로 검증하기 위해 단일마커(single-marker)
    및 다중마커(multi-marker) 분석을 수행하였다. 단일마커 분석은 각
    균을 독립적으로 평가하여 차등 풍부도(differential abundance)를
    검정하였으며(edgeR, ANCOM 등), 다중마커 분석은 여러 균을 동시에
    분석하여 변수 선택(feature selection)을 수행하였다(LASSO, gLASSO 등). 검출된 균의 생물학적 타당성은 OHMI 데이터베이스를
    활용한 문헌 검증 및 집합군 풍부도 분석(Taxon Set Enrichment
    Analysis, TSEA)를 통해 확인하였다.
    본 시뮬레이션 연구는 희소성(sparsity)과 효과 크기(effect
    size)의 수준을 달리하며 HisCoM-microb 의 성능을 평가하였다. 효과
    크기가 증가함에 따라 HisCoM-microb 의 장점은 더욱 뚜렷하게
    나타났으며, 특히 낮거나 높은 희소성 조건에서 우수한 성능을 보였다.
    실제 CRC 데이터 분석에서 HisCoM-microb 은 비교된 방법 중 높은
    참발견률(True Discovery Rate, TDR)을 달성하였다. 이는 질병과
    실제로 연관된 마커를 효과적으로 식별할 수 있음을 시사한다. 또한
    TSEA 를 통해 검출된 균의 생물학적 연관성을 추가로 확인하였다. 본
    모델은 대표적인 CRC 연관 계통으로 알려진 Fusobacteria 계통을
    식별하였으며, 기존 방법에서 탐지되지 않았던 Coprobacillus 및
    Coprococcus 와 같은 CRC 연관 균도 추가로 검출하였다.
    HisCoM-microb 은 OTU 와 이에 대응하는 상위 분류군을 함께
    고려하는 2 단계(two-layer) 모델링 구조를 구현하여 생물학적 정보를
    모델 내에 반영한다. 본 연구는 CRC 데이터를 통해 그 실용성과
    생물학적 타당성을 입증하였다. 향후 다양한 질병 데이터를 통해 본
    모델을 추가 검증함으로써, 바이오마커 발굴(biomarker discovery) 및
    실제 연구와 진단에 활용될 수 있을 것으로 기대된다.
    번역하기

    마이크로바이옴 연구에서 특정 미생물의 풍부도(abundance) 변화를 기반으로 한 접근법은 질병 연관성을 분석하는 데 널리 사용되고 있다. 한편, 실제 마이크로바이옴 데이터는 희소성(sparsity),...

    마이크로바이옴 연구에서 특정 미생물의 풍부도(abundance)
    변화를 기반으로 한 접근법은 질병 연관성을 분석하는 데 널리
    사용되고 있다. 한편, 실제 마이크로바이옴 데이터는 희소성(sparsity),
    구성비 제약성(compositionality), 고차원성(high dimensionality),
    그리고 분류학적 계층 구조(taxonomic hierarchy)와 같은 고유한
    특성을 가진다. 그러나 이러한 특성들은 기존의 통계적 방법들(e.g.,
    edgeR, metagenomeSeq, ANCOM 등)에서 충분히 반영되지 못한다.
    또한 일부 기존 방법은 다른 균(taxa)과 상관성으로 인해 제외하는 등
    생물학적으로 중요한 균(taxa)의 발견을 놓칠 수 있다. 그 결과,
    이러한 접근법들은 마이크로바이옴 데이터에 적용한 기존 통계 모델의
    한계로 인해 생물학적 해석 가능성이 제한되는 문제를 갖는다.
    이러한 한계를 보완하기 위해 본 연구는 마이크로바이옴 데이터를
    위한 계층적 구조 성분 모델(hierarchical structured component
    model incorporating taxonomic hierarchy for microbiome data,
    HisCoM-microb)을 제안한다. 본 모델은 다음 세 가지 방식으로
    마이크로바이옴 데이터의 분석적 문제를 해결하도록 설계되었다: (1)
    순열 기반(permutation-based) 접근법을 통해 희소성 문제를
    완화하고, (2) 중심화 로그비(centered log-ratio, CLR) 변환을
    사용하여 구성비 제약을 보정하며, (3) 분류학적 계층 정보를 모델
    구조에 내재함으로써 생물학적 정보를 반영한다.
    본 연구는 HisCoM-microb 을 시뮬레이션 데이터와 실제
    대장암(colorectal cancer, CRC) 마이크로바이옴 데이터(120 명의
    CRC 환자와 172 명의 대조군, 총 335 OTU)를 이용해 평가하였다.
    모델의 성능을 종합적으로 검증하기 위해 단일마커(single-marker)
    및 다중마커(multi-marker) 분석을 수행하였다. 단일마커 분석은 각
    균을 독립적으로 평가하여 차등 풍부도(differential abundance)를
    검정하였으며(edgeR, ANCOM 등), 다중마커 분석은 여러 균을 동시에
    분석하여 변수 선택(feature selection)을 수행하였다(LASSO, gLASSO 등). 검출된 균의 생물학적 타당성은 OHMI 데이터베이스를
    활용한 문헌 검증 및 집합군 풍부도 분석(Taxon Set Enrichment
    Analysis, TSEA)를 통해 확인하였다.
    본 시뮬레이션 연구는 희소성(sparsity)과 효과 크기(effect
    size)의 수준을 달리하며 HisCoM-microb 의 성능을 평가하였다. 효과
    크기가 증가함에 따라 HisCoM-microb 의 장점은 더욱 뚜렷하게
    나타났으며, 특히 낮거나 높은 희소성 조건에서 우수한 성능을 보였다.
    실제 CRC 데이터 분석에서 HisCoM-microb 은 비교된 방법 중 높은
    참발견률(True Discovery Rate, TDR)을 달성하였다. 이는 질병과
    실제로 연관된 마커를 효과적으로 식별할 수 있음을 시사한다. 또한
    TSEA 를 통해 검출된 균의 생물학적 연관성을 추가로 확인하였다. 본
    모델은 대표적인 CRC 연관 계통으로 알려진 Fusobacteria 계통을
    식별하였으며, 기존 방법에서 탐지되지 않았던 Coprobacillus 및
    Coprococcus 와 같은 CRC 연관 균도 추가로 검출하였다.
    HisCoM-microb 은 OTU 와 이에 대응하는 상위 분류군을 함께
    고려하는 2 단계(two-layer) 모델링 구조를 구현하여 생물학적 정보를
    모델 내에 반영한다. 본 연구는 CRC 데이터를 통해 그 실용성과
    생물학적 타당성을 입증하였다. 향후 다양한 질병 데이터를 통해 본
    모델을 추가 검증함으로써, 바이오마커 발굴(biomarker discovery) 및
    실제 연구와 진단에 활용될 수 있을 것으로 기대된다.

    더보기

    목차 (Table of Contents)

    • 1 INTRODUCTION 1
    • 1.1. Background of Huamn Microbiome Research 1
    • 1.2. Analytical Characteristics of Microbiome Data 3
    • 1.2.1. Sparsity 3
    • 1.2.2. Compositionality 3
    • 1 INTRODUCTION 1
    • 1.1. Background of Huamn Microbiome Research 1
    • 1.2. Analytical Characteristics of Microbiome Data 3
    • 1.2.1. Sparsity 3
    • 1.2.2. Compositionality 3
    • 1.2.3. High-Dimensionality 4
    • 1.2.4. Taxonomic Hierarchy 4
    • 1.3. Existing Analytical Approaches 6
    • 1.3.1. Single-Marker 6
    • 1.3.2. Multi-Marker 7
    • 1.4. Limitation of Conventional Methods 9
    • 1.5. Research Objectives 10
    • 2 METHODS AND MATERIALS 11
    • 2.1. Hierarchical Structured Component Model for Microbiome Data (HisCoM-microb) 11
    • 2.1.1. Data Transformation: Centered Log Ratio Transformation 13
    • 2.1.2. Model Structure: OTU-Taxon Hierarchy and Taxon Phenotype Association 14
    • 2.1.3. Estimation 15
    • 2.1.4. Statistical Inference via Permutation Testing 17
    • 2.2. Comparison Methods 18
    • 2.2.1. Single-Marker 18
    • 2.2.2. Multi-Marker 21
    • 2.3. Disease Relevance Analysis using Taxon Set Enrichment Analysis (TSEA) 23
    • 2.4. Materials: Colorectal Cancer (CRC) Microbiome Data 24
    • 2.4.1. Data Acquisition 24
    • 2.4.2. Filtering and Quality Control 24
    • 2.4.3. Final Dataset 25
    • 3 SIMULATION STUDY 28
    • 3.1. Simulation Design 28
    • 3.1.1. Construction of Simulated Microbial Environment 30
    • 3.1.2. Causal Taxon Selection 32
    • 3.1.3. Simulation Model 32
    • 3.1.4. Method Comparison with Simulation Model 33
    • 3.1.5. Performance Evaluation 33
    • 3.2. Results: Single-Marker Scenarios 35
    • 3.2.1. Type I Error Rate 35
    • 3.2.2. Statistical Power 38
    • 3.3. Results: Multi-Marker Scenario 42
    • 3.3.1. Type I Error Rate 42
    • 3.3.2. Statistical Power 44
    • 4 COLORECTAL CANCER MICROBIOME ANALYSIS 45
    • 4.1. Exploratory Data Analysis 45
    • 4.1.1. Taxonomic Composition 45
    • 4.1.2. Alpha Diversity and Beta Diversity 48
    • 4.2. Single-Marker Analysis 50
    • 4.2.1. Comparison of Identified Markers across Taxonomic Levels 50
    • 4.2.2. Identified Microbial Markers and Literature Review 51
    • 4.3. Multi-Marker Analysis 54
    • 4.3.1. True Discovery Rate (TDR)–based Validation 54
    • 4.3.2. AUC-based Evaluation of Predictive Performance 57
    • 4.4. Significant Taxa and OTUs Identified by HisCoM-microb in Taxonomy Hierarchy 59
    • 4.4.1. Consistent and Inconsistent Relationships Observed in Taxonomic Hierarchy 59
    • 4.4.2. Identification and Validation of Significant OTUs 59
    • 4.4.3. Associations of Significant OTUs with CRC 60
    • 4.5. Taxon Set Enrichment Analysis (TSEA) with HisCoM-microb-Identified Taxa 63
    • 4.5.1. CRC-Associated Taxon Sets 63
    • 4.5.2. Secondary-Associated Taxon Sets 63
    • 5 DISCUSSION 68
    • 5.1. Incorporating Taxonomic Hierarchy as Biological Context in HisCoM-microb 68
    • 5.2. Performance under Different Levels of Sparsity and Effect Size 69
    • 5.3. Identification of Biologically Relevant Markers in Real CRC Data 70
    • 5.4. Validation and Disease-Related Interpretation of Identified Taxa 71
    • 5.5. Application of HisCoM-microb in Diverse Analytical Frameworks 72
    • 5.6. Computational Efficiency Analysis and Comparison 74
    • 5.7. Limitations and Future Work 77
    • 6 CONCLUSION 78
    • Bibliography 80
    • Abstract in Korean 89
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼