RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    From Proteins, to Machines, to Protons, to Genes, and Back Again.

    한글로보기

    https://www.riss.kr/link?id=T16616236

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The success of data standards and public databases in biology is the foundation for the current and continued success of machine learning in biology and medicine. This dissertation explores the interactions between biology, computers, and people in order to develop novel machine learning methods to model complex biological problems. Data is one of the main resources to do machine learning, and Chapters 1, 2, 3 are explicitly about data organization and quality assurance in the protein Nuclear Magnetic Resonance (NMR) spectroscopy discipline. Chapters 4 and 5 present new machine learning architectures to address learning tasks in genomic site recognition and NMR chemical shift prediction. Chapter 1 investigates the manner protein NMR chemical shift data is deposited at the Biological Magnetic Resonance Bank (BMRB) in order to build simple table look-up models to estimate protein chemical shifts. In Chapter 1, we find there is low sequence diversity and data redundancy in the BMRB that was a challenge to locate and filter out. Without filtering out BMRB entries with the same sequence, and possibly the same chemical shifts, look-up models will be more accurate due to data contamination in training and testing sets. Chapter 2 examines approaches to curate a large protein sample production and NMR database to create an NMR time-domain dataset. Quality assurance tests in this NMR sample/FID database uncovered data collisions and redundancies among the database records, which motivated the development of new NMR database management tools. Chapter 3 presents a relational database schema to archive protein NMR samples and associated time-domain data called SpecDB. SpecDB is open source and available at https://github.rpi.edu/RPIBioinformatics/SpecDB.git. Chapter 4 explores how deep neural networks can recognize genomic splice acceptor and donor sites from sequence alone, achieving 97% accuracy for highly used splice donor sites. Chapter 4 also investigates neural networks for intron/exon sequence classification, maximally reaching 77% accuracy. Chapter 5 presents the application of marginalized graph kernels to prediction of NMR chemical shifts for small organic molecules. Incorporating chemical descriptors to graph kernels reaches a 3.501 ppm mean absolute error for Carbon chemical shifts. In total, the following five dissertation chapters explore work in data integrity, organization, and learning techniques from data for applications to structural biology problems.
    번역하기

    The success of data standards and public databases in biology is the foundation for the current and continued success of machine learning in biology and medicine. This dissertation explores the interactions between biology, computers, and people in o...

    The success of data standards and public databases in biology is the foundation for the current and continued success of machine learning in biology and medicine. This dissertation explores the interactions between biology, computers, and people in order to develop novel machine learning methods to model complex biological problems. Data is one of the main resources to do machine learning, and Chapters 1, 2, 3 are explicitly about data organization and quality assurance in the protein Nuclear Magnetic Resonance (NMR) spectroscopy discipline. Chapters 4 and 5 present new machine learning architectures to address learning tasks in genomic site recognition and NMR chemical shift prediction. Chapter 1 investigates the manner protein NMR chemical shift data is deposited at the Biological Magnetic Resonance Bank (BMRB) in order to build simple table look-up models to estimate protein chemical shifts. In Chapter 1, we find there is low sequence diversity and data redundancy in the BMRB that was a challenge to locate and filter out. Without filtering out BMRB entries with the same sequence, and possibly the same chemical shifts, look-up models will be more accurate due to data contamination in training and testing sets. Chapter 2 examines approaches to curate a large protein sample production and NMR database to create an NMR time-domain dataset. Quality assurance tests in this NMR sample/FID database uncovered data collisions and redundancies among the database records, which motivated the development of new NMR database management tools. Chapter 3 presents a relational database schema to archive protein NMR samples and associated time-domain data called SpecDB. SpecDB is open source and available at https://github.rpi.edu/RPIBioinformatics/SpecDB.git. Chapter 4 explores how deep neural networks can recognize genomic splice acceptor and donor sites from sequence alone, achieving 97% accuracy for highly used splice donor sites. Chapter 4 also investigates neural networks for intron/exon sequence classification, maximally reaching 77% accuracy. Chapter 5 presents the application of marginalized graph kernels to prediction of NMR chemical shifts for small organic molecules. Incorporating chemical descriptors to graph kernels reaches a 3.501 ppm mean absolute error for Carbon chemical shifts. In total, the following five dissertation chapters explore work in data integrity, organization, and learning techniques from data for applications to structural biology problems.

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼