RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Small size sampling of large size metagenomic sequences and its environmental applications

    한글로보기

    https://www.riss.kr/link?id=T14858935

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Metagenomics is regarded as one of the great opportunities of environmental biotechnology, as high throughput DNA sequencing is getting popular. However, it requires a large amount of computing power and heavy time consumption for analyzing huge size of DNA sequence data, which is main barrier to the applicability of metagenomics in environmental applications. To shorten computation time with less computing resources, a simple concept of rarefaction technique was utilized in this study. The research objectives of this study were to examine the accuracy of the taxonomical and functional analyses by the small sampling method optimized in this work, and to provide specific guideline information on the size and the way of small sampling, and the levels of taxonomical and functional information.
    As the first step, metagenomic sequences of a mock community, which was intentionally made of known bacterial strains to get a mixed genome, were examined with BLAST search. In total 792 samples were made with 6 different sampling methods. Since the real taxonomic information of the mock community is known, one can evaluate how well a sample represents the original taxonomic information. As the second step, the sampling method was also tested for 10 known results of previously reported metagenome researches, comparing taxonomic information, as well as gene function information.
    The BLAST search test showed that the increase of the sample size results the increase of accuracy in taxonomic distribution, following the statistical tendency. Assuming a general normal distribution model, a sample with 5,000 reads was suggested as an option with 1% margin of error and 85% confidence. Meanwhile, among the 6 sampling methods, the convenience sampling method of selecting from the end showed the noticeably worst representativeness. While other method showed much better results than that, the systematic sampling of uniform selection showed the best result.
    Among the results from previously reported metagenome sequence results, 9 out of 10 cases of both the most annotated phyla and the most annotated classes were matched between the samples and the originals. The Shannon diversity indexes calculated from them also showed the similarity between ones from the samples and ones from the originals in their tendency. Additionally, COG and SEED Subsystems function types annotated by MG-RAST showed the similarity. From the findings, a systematic method for determining size of credible sample was proposed with its users’ guideline including a flow chart for designing the small sampling method and evaluating its errors.
    Since this sampling method only costs short time and small computing resources, one can use this approach to develop a standard or a protocol to preview or pre-check metagenomics data, before performing more accurate analysis with original full sequences. This will make metagenomics to be more efficient, enabling its environmental applications in wider area.
    번역하기

    Metagenomics is regarded as one of the great opportunities of environmental biotechnology, as high throughput DNA sequencing is getting popular. However, it requires a large amount of computing power and heavy time consumption for analyzing huge size ...

    Metagenomics is regarded as one of the great opportunities of environmental biotechnology, as high throughput DNA sequencing is getting popular. However, it requires a large amount of computing power and heavy time consumption for analyzing huge size of DNA sequence data, which is main barrier to the applicability of metagenomics in environmental applications. To shorten computation time with less computing resources, a simple concept of rarefaction technique was utilized in this study. The research objectives of this study were to examine the accuracy of the taxonomical and functional analyses by the small sampling method optimized in this work, and to provide specific guideline information on the size and the way of small sampling, and the levels of taxonomical and functional information.
    As the first step, metagenomic sequences of a mock community, which was intentionally made of known bacterial strains to get a mixed genome, were examined with BLAST search. In total 792 samples were made with 6 different sampling methods. Since the real taxonomic information of the mock community is known, one can evaluate how well a sample represents the original taxonomic information. As the second step, the sampling method was also tested for 10 known results of previously reported metagenome researches, comparing taxonomic information, as well as gene function information.
    The BLAST search test showed that the increase of the sample size results the increase of accuracy in taxonomic distribution, following the statistical tendency. Assuming a general normal distribution model, a sample with 5,000 reads was suggested as an option with 1% margin of error and 85% confidence. Meanwhile, among the 6 sampling methods, the convenience sampling method of selecting from the end showed the noticeably worst representativeness. While other method showed much better results than that, the systematic sampling of uniform selection showed the best result.
    Among the results from previously reported metagenome sequence results, 9 out of 10 cases of both the most annotated phyla and the most annotated classes were matched between the samples and the originals. The Shannon diversity indexes calculated from them also showed the similarity between ones from the samples and ones from the originals in their tendency. Additionally, COG and SEED Subsystems function types annotated by MG-RAST showed the similarity. From the findings, a systematic method for determining size of credible sample was proposed with its users’ guideline including a flow chart for designing the small sampling method and evaluating its errors.
    Since this sampling method only costs short time and small computing resources, one can use this approach to develop a standard or a protocol to preview or pre-check metagenomics data, before performing more accurate analysis with original full sequences. This will make metagenomics to be more efficient, enabling its environmental applications in wider area.

    더보기

    참고문헌 (Reference)

    1. 농업용수 수질개선사업, 어디까지 왔나? PRI Focus 20, Kim, H., ISSN:2288-0682, , 2013

    1. 농업용수 수질개선사업, 어디까지 왔나? PRI Focus 20, Kim, H., ISSN:2288-0682, , 2013

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼