RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    단백질 합성효소(ARS)의 세포 조절 기작에 관한 가설 생성 연구

    한글로보기

    https://www.riss.kr/link?id=T15875687

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Conceptual biology is a new research trend in the field of biomedical science that seeks to gain
    new insights by collecting fragmented knowledge by reviewing the findings of biomedical fields
    accumulated in multiple databases and linking related concepts. The concept of conceptual biology
    has also been applied to modern drug developments, with the rapid development of text mining
    techniques to extract desired information from biomedical data, such as paper publications,
    systematically well managed.

    However, conceptual biology is still not the main methodology for biomedical and new drug
    development research. This is because there is a deep-rooted perception among researchers that
    hypotheses generated and proposed in an automated manner through biomedical text mining
    technology are less reliable than those made by conventional research methods. Therefore, prior to
    recklessly introducing new hypotheses and drug targets under the banner of conceptual biology, it
    is necessary to show that the research results obtained through the existing experimental-oriented
    research methods can be derived equally by computerized text mining techniques.

    Therefore, in this work, we first demonstrate how reliable the connections between entities
    extracted from unstructured text data are through comparisons with existing research cases, and
    then apply the methodology to actual new drug target discovery to propose new hypotheses and
    materials to researchers. In this process, we used the theoretical framework of literature-based
    discovery (LBD) to bridge between conceptual biology and the field of library and information
    science, and used an enzyme called aminoacyl-tRNA synthetase (ARS), which has recently been
    spotlighted as a criterion for finding drug target proteins, as a key research material.

    To implement the above intentions and configurations in practice, the research procedure was
    also conducted in two major steps. First, in the first part, we summarized the existing research
    results related to ARS and expressed them in the form of a series of paths, and then, examined
    whether they could be reproduced in the same way through the currently used text mining
    technique. In this process, research papers related to the three pairs of ARS and amino acids
    (LARS1/leucine, QARS1/glutamine, MARS1/methionine) were used as the standard. And by
    comparing the results when the literature related only to the contents of LARS1/leucine,
    QARS1/glutamine, and MARS1/methionine were used separately with when all the literatures
    were combined, we wanted to visually identify the benefits of increasing the size of the literature
    group to be analyzed.

    As a result of the experiments, it was confirmed that most of the contents of the standard papers
    were reproduced through text mining techniques, demonstrating that this research method is fully
    utilizable unlike researchers' stereotypes. In particular, it is shown that the papers are better
    represented when using the entire data compared to utilizing only the literature groups directly
    related to the hypothesis we want to generate, indicating that even if we aim to generate
    hypotheses in a particular ARS field, it is better to utilize the data around it. Also, throughout this
    process, we also coordinate and optimize the specifications of the methodology, including the type
    of named entity dictionary to be used, the stop words list to be included, and the categories of data
    to be analyzed, to lay the foundation for what will follow.

    Meanwhile, in the latter part of the study, we explored substances that mediated the interaction
    between ARS and amino acids, and proposed new features of ARS which have not yet been
    proven. Especially, ranking algorithm was applied to the derived paths, allowing researchers to
    preferentially review materials that are expected to be worth studying as targets for new drug
    development. To this end, a total of three ranking mechanisms were devised and utilized: using the
    centrality indicators of words, considering the frequency of relations, and applying the ratio of
    relation frequency. And, as a result, in case of the second method, utilizing the frequency of
    relations, the results of previous studies in the first half were found to be at the highest rank.

    In addition, the last one section of the latter part of the study was assigned to the contents of
    substituting our methodology to the field of WARS1/tryptophan, where all operating mechanisms
    were not fully elucidated, being added a new attempt to express in an explicit path what has not
    yet been directly identified. Through this, we reinforced the usefulness of our research
    methodology by reconfirming the contents suggested only as hypotheses through partial
    experiments and contexts. At the same time, we were also able to enjoy the effect of presenting
    other possible paths that could lead to the development of new drugs.

    Despite the existence of some limitations, such as limited use of relation classification models
    using deep learning and incompleteness of text subject to analysis solely on titles and abstracts,
    this work is significant in that it systematically proves that the methodology of conceptual biology
    can be used in drug target discovery on several grounds. Starting with this study, if conceptual
    biology is more commonly used for biomedical and drug development research, it is expected that
    various social and economic costs incurred in the related research process will be dramatically
    reduced.
    번역하기

    Conceptual biology is a new research trend in the field of biomedical science that seeks to gain new insights by collecting fragmented knowledge by reviewing the findings of biomedical fields accumulated in multiple databases and linking related conce...

    Conceptual biology is a new research trend in the field of biomedical science that seeks to gain
    new insights by collecting fragmented knowledge by reviewing the findings of biomedical fields
    accumulated in multiple databases and linking related concepts. The concept of conceptual biology
    has also been applied to modern drug developments, with the rapid development of text mining
    techniques to extract desired information from biomedical data, such as paper publications,
    systematically well managed.

    However, conceptual biology is still not the main methodology for biomedical and new drug
    development research. This is because there is a deep-rooted perception among researchers that
    hypotheses generated and proposed in an automated manner through biomedical text mining
    technology are less reliable than those made by conventional research methods. Therefore, prior to
    recklessly introducing new hypotheses and drug targets under the banner of conceptual biology, it
    is necessary to show that the research results obtained through the existing experimental-oriented
    research methods can be derived equally by computerized text mining techniques.

    Therefore, in this work, we first demonstrate how reliable the connections between entities
    extracted from unstructured text data are through comparisons with existing research cases, and
    then apply the methodology to actual new drug target discovery to propose new hypotheses and
    materials to researchers. In this process, we used the theoretical framework of literature-based
    discovery (LBD) to bridge between conceptual biology and the field of library and information
    science, and used an enzyme called aminoacyl-tRNA synthetase (ARS), which has recently been
    spotlighted as a criterion for finding drug target proteins, as a key research material.

    To implement the above intentions and configurations in practice, the research procedure was
    also conducted in two major steps. First, in the first part, we summarized the existing research
    results related to ARS and expressed them in the form of a series of paths, and then, examined
    whether they could be reproduced in the same way through the currently used text mining
    technique. In this process, research papers related to the three pairs of ARS and amino acids
    (LARS1/leucine, QARS1/glutamine, MARS1/methionine) were used as the standard. And by
    comparing the results when the literature related only to the contents of LARS1/leucine,
    QARS1/glutamine, and MARS1/methionine were used separately with when all the literatures
    were combined, we wanted to visually identify the benefits of increasing the size of the literature
    group to be analyzed.

    As a result of the experiments, it was confirmed that most of the contents of the standard papers
    were reproduced through text mining techniques, demonstrating that this research method is fully
    utilizable unlike researchers' stereotypes. In particular, it is shown that the papers are better
    represented when using the entire data compared to utilizing only the literature groups directly
    related to the hypothesis we want to generate, indicating that even if we aim to generate
    hypotheses in a particular ARS field, it is better to utilize the data around it. Also, throughout this
    process, we also coordinate and optimize the specifications of the methodology, including the type
    of named entity dictionary to be used, the stop words list to be included, and the categories of data
    to be analyzed, to lay the foundation for what will follow.

    Meanwhile, in the latter part of the study, we explored substances that mediated the interaction
    between ARS and amino acids, and proposed new features of ARS which have not yet been
    proven. Especially, ranking algorithm was applied to the derived paths, allowing researchers to
    preferentially review materials that are expected to be worth studying as targets for new drug
    development. To this end, a total of three ranking mechanisms were devised and utilized: using the
    centrality indicators of words, considering the frequency of relations, and applying the ratio of
    relation frequency. And, as a result, in case of the second method, utilizing the frequency of
    relations, the results of previous studies in the first half were found to be at the highest rank.

    In addition, the last one section of the latter part of the study was assigned to the contents of
    substituting our methodology to the field of WARS1/tryptophan, where all operating mechanisms
    were not fully elucidated, being added a new attempt to express in an explicit path what has not
    yet been directly identified. Through this, we reinforced the usefulness of our research
    methodology by reconfirming the contents suggested only as hypotheses through partial
    experiments and contexts. At the same time, we were also able to enjoy the effect of presenting
    other possible paths that could lead to the development of new drugs.

    Despite the existence of some limitations, such as limited use of relation classification models
    using deep learning and incompleteness of text subject to analysis solely on titles and abstracts,
    this work is significant in that it systematically proves that the methodology of conceptual biology
    can be used in drug target discovery on several grounds. Starting with this study, if conceptual
    biology is more commonly used for biomedical and drug development research, it is expected that
    various social and economic costs incurred in the related research process will be dramatically
    reduced.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    Conceptual biology는 여러 데이터베이스에 축적된 생물 의학 분야의 연구 결과물들을 개념 중심으로 검토하고 관련된 개념들을 서로 연결 지음으로써, 단편적으로 산재해 있는 지식들을 모아 새로운 통찰을 얻고자 하는 생물 의학 분야의 새로운 연구 조류이다. 이러한 conceptual biology의 개념은 현대의 신약 개발 과정에도 응용되고 있는데, 타 분야에 비해 논문 출판물과 같은 비정형 데이터가 체계적으로 잘 관리되고 있는 생물 의학의 분야적 특성과 더불어 이들 데이터로부터 원하는 정보를 추출 해낼 수 있는 텍스트 마이닝 기법이 빠르게 발전하고 있다는 점 또한 그 원인이 되고 있다.

    하지만 여전히 conceptual biology는 생물 의학 및 신약 개발 연구에 있어 주된 방법론으로 자리하지는 못하고 있다. 이는 바이오 텍스트 마이닝 기술을 통해 자동화된 방식으로 생성되고 제안된 가설들이 기존의 연구 방식에 의한 그것들에 비해 덜 믿을만하다는 인식이 연구자들 사이에 뿌리 깊게 퍼져 있기 때문이다. 따라서 무작정 conceptual biology의 기치 아래 새로운 가설 및 약물 표적을 내놓기에 앞서, 기존의 실험 중심 연구 방식으로 얻어낸 연구 결과들을 컴퓨터를 활용한 텍스트 마이닝 기법을 통해서 역시 동일하게 도출해낼 수 있음을 보일 필요성이 있다.

    따라서 본 연구에서는 텍스트 마이닝 기법을 통해 비정형 텍스트로부터 추출된 개체 간 연결 관계들이 얼마나 믿을만한 지를 기존의 연구 사례들과의 비교를 통해 먼저 입증한 후, 해당 방법론을 실질적인 신약 표적 발굴에 적용하여 새로운 가설 및 물질들을 관련 연구자들에게 제안하는 방식을 택했다. 이 과정에서 문헌 기반 발견법이라는 이론적 틀을 활용하여 conceptual biology와 문헌정보학 분야 간의 다리를 놓았으며, 최근 약물 표적 단백질을 찾아내기 위한 기준으로서 각광받고 있는 ARS (aminoacyl-tRNA synthetase)라는 효소를 핵심 연구 소재로 삼았다.

    상기한 의도와 구성을 실제 구현하기 위해, 연구 절차 또한 크게 두 단계로 나뉘어 수행되었다. 먼저, 전반부에서는 ARS와 관련된 기존의 연구 결과들을 요약하여 일련의 경로 형태로 표현하고, 현재 사용하고 있는 텍스트 마이닝 기법을 통해 이들을 똑같이 재현해 낼 수 있는지를 살폈다. 이 과정에서 세 쌍의 ARS 및 아미노산 LARS1/leucine, QARS1/glutamine, MARS1/methionine 등과 관련된 세 부류의 연구 논문을 기준으로 하였으며, LARS1/leucine, QARS1/glutamine, MARS1/methionine 각각의 내용과만 관련된 문헌을 별도로 활용한 경우와 전체 문헌을 모두 합쳤을 시의 결과를 비교함으로써 분석 대상 문헌 집단의 크기를 늘리는 데에서 오는 이점을 가시적으로 확인하고자 하였다.

    실험 결과, 기준이 되었던 논문의 내용들이 텍스트 마이닝 기법을 통해서도 대부분 재현되는 것이 확인되어, 연구자들의 고정 관념과는 달리 본 연구 방식이 충분히 활용할 만한 것임을 입증할 수 있었다. 특히, 생성하고자 하는 가설과 직접적으로 관련된 문헌 집단만을 활용하는 것에 비해 수집한 전체 데이터를 이용할 때에 논문의 내용이 보다 잘 재현되는 것으로 나타나, 특정 ARS 분야의 가설을 생성하는 것을 목표로 하더라도 그 주변부의 데이터까지 활용하는 것이 옳은 판단임을 알 수 있었다. 또한, 이러한 과정을 거치며 사용해야 할 개체명 사전의 종류, 불용어로 포함시킬 단어의 목록, 이용할 데이터의 범주 등 해당 방법론의 구체적인 사항들을 조정하고 최적화하여, 이어질 내용을 위한 기반을 마련하였다.

    한편, 연구의 후반부에서는 ARS와 아미노산 간의 상호작용을 매개하는 물질들을 발굴하고, 이를 통해 아직까지 증명되지 않은 ARS의 새로운 기능들을 제안하는 내용을 다루었다. 특히, 도출된 경로 집단에 순위화 방법론을 적용하여, 신약 개발의 표적으로서 연구할 가치가 있을 것으로 예상되는 물질 및 이를 포함한 경로들을 우선적으로 검토할 수 있게끔 하였다. 이를 위해 단어의 중심성 지표를 활용하는 방법, 관계의 빈도를 이용하는 방식, 관계 빈도의 비율을 적용하는 방안 등 총 세 가지 순위화 기제를 고안하여 활용하였는데, 관계의 빈도를 활용했을 때에 전반부에서 이용되었던 종전의 연구 결과들이 가장 높은 순위에 위치함을 확인할 수 있었다.

    또한, 후반부 연구의 마지막 한 절을 모든 작동 기제가 완전하게 규명되지 않은 WARS1/tryptophan 관련 분야에 본 연구의 방법론을 대입하는 내용에 할당하여, 아직 직접적으로 밝혀지지는 않았지만 지금까지의 연구를 통해 유추만 가능했던 내용을 명시적인 경로로 표현해 보는 새로운 시도를 덧붙였다. 이를 통해, 부분적인 실험과 정황을 통해서 가설로만 제안되었던 내용을 문헌 기반 발견법을 통해 다시 한번 확인함으로써 본 연구 방법론의 유용성을 강화함과 동시에, 신약 개발로 이어질 가능성이 있는 또 다른 경로들을 함께 제시하는 효과를 누릴 수 있었다.

    본 연구는 딥러닝을 활용한 관계 분류 모델이 제한적으로만 활용되었다는 점, 분석 대상 텍스트가 제목 및 초록에 한정되어 있다는 점 등 몇몇 한계점이 존재함에도 불구하고, conceptual biology의 방법론이 신약 표적 물질 발굴에 실질적으로 이용될 수 있음을 여러 근거를 들어 체계적으로 증명했다는 점에서 큰 의의를 지닌다고 볼 수 있다. 본 연구를 기점으로 conceptual biology 방법론이 생물 의학 및 신약 개발 연구에 더욱 보편적으로 활용된다면, 관련 연구 과정상에 발생하는 각종 사회적·경제적 비용을 획기적으로 절감할 수 있을 것으로 기대한다.
    번역하기

    Conceptual biology는 여러 데이터베이스에 축적된 생물 의학 분야의 연구 결과물들을 개념 중심으로 검토하고 관련된 개념들을 서로 연결 지음으로써, 단편적으로 산재해 있는 지식들을 모아 새...

    Conceptual biology는 여러 데이터베이스에 축적된 생물 의학 분야의 연구 결과물들을 개념 중심으로 검토하고 관련된 개념들을 서로 연결 지음으로써, 단편적으로 산재해 있는 지식들을 모아 새로운 통찰을 얻고자 하는 생물 의학 분야의 새로운 연구 조류이다. 이러한 conceptual biology의 개념은 현대의 신약 개발 과정에도 응용되고 있는데, 타 분야에 비해 논문 출판물과 같은 비정형 데이터가 체계적으로 잘 관리되고 있는 생물 의학의 분야적 특성과 더불어 이들 데이터로부터 원하는 정보를 추출 해낼 수 있는 텍스트 마이닝 기법이 빠르게 발전하고 있다는 점 또한 그 원인이 되고 있다.

    하지만 여전히 conceptual biology는 생물 의학 및 신약 개발 연구에 있어 주된 방법론으로 자리하지는 못하고 있다. 이는 바이오 텍스트 마이닝 기술을 통해 자동화된 방식으로 생성되고 제안된 가설들이 기존의 연구 방식에 의한 그것들에 비해 덜 믿을만하다는 인식이 연구자들 사이에 뿌리 깊게 퍼져 있기 때문이다. 따라서 무작정 conceptual biology의 기치 아래 새로운 가설 및 약물 표적을 내놓기에 앞서, 기존의 실험 중심 연구 방식으로 얻어낸 연구 결과들을 컴퓨터를 활용한 텍스트 마이닝 기법을 통해서 역시 동일하게 도출해낼 수 있음을 보일 필요성이 있다.

    따라서 본 연구에서는 텍스트 마이닝 기법을 통해 비정형 텍스트로부터 추출된 개체 간 연결 관계들이 얼마나 믿을만한 지를 기존의 연구 사례들과의 비교를 통해 먼저 입증한 후, 해당 방법론을 실질적인 신약 표적 발굴에 적용하여 새로운 가설 및 물질들을 관련 연구자들에게 제안하는 방식을 택했다. 이 과정에서 문헌 기반 발견법이라는 이론적 틀을 활용하여 conceptual biology와 문헌정보학 분야 간의 다리를 놓았으며, 최근 약물 표적 단백질을 찾아내기 위한 기준으로서 각광받고 있는 ARS (aminoacyl-tRNA synthetase)라는 효소를 핵심 연구 소재로 삼았다.

    상기한 의도와 구성을 실제 구현하기 위해, 연구 절차 또한 크게 두 단계로 나뉘어 수행되었다. 먼저, 전반부에서는 ARS와 관련된 기존의 연구 결과들을 요약하여 일련의 경로 형태로 표현하고, 현재 사용하고 있는 텍스트 마이닝 기법을 통해 이들을 똑같이 재현해 낼 수 있는지를 살폈다. 이 과정에서 세 쌍의 ARS 및 아미노산 LARS1/leucine, QARS1/glutamine, MARS1/methionine 등과 관련된 세 부류의 연구 논문을 기준으로 하였으며, LARS1/leucine, QARS1/glutamine, MARS1/methionine 각각의 내용과만 관련된 문헌을 별도로 활용한 경우와 전체 문헌을 모두 합쳤을 시의 결과를 비교함으로써 분석 대상 문헌 집단의 크기를 늘리는 데에서 오는 이점을 가시적으로 확인하고자 하였다.

    실험 결과, 기준이 되었던 논문의 내용들이 텍스트 마이닝 기법을 통해서도 대부분 재현되는 것이 확인되어, 연구자들의 고정 관념과는 달리 본 연구 방식이 충분히 활용할 만한 것임을 입증할 수 있었다. 특히, 생성하고자 하는 가설과 직접적으로 관련된 문헌 집단만을 활용하는 것에 비해 수집한 전체 데이터를 이용할 때에 논문의 내용이 보다 잘 재현되는 것으로 나타나, 특정 ARS 분야의 가설을 생성하는 것을 목표로 하더라도 그 주변부의 데이터까지 활용하는 것이 옳은 판단임을 알 수 있었다. 또한, 이러한 과정을 거치며 사용해야 할 개체명 사전의 종류, 불용어로 포함시킬 단어의 목록, 이용할 데이터의 범주 등 해당 방법론의 구체적인 사항들을 조정하고 최적화하여, 이어질 내용을 위한 기반을 마련하였다.

    한편, 연구의 후반부에서는 ARS와 아미노산 간의 상호작용을 매개하는 물질들을 발굴하고, 이를 통해 아직까지 증명되지 않은 ARS의 새로운 기능들을 제안하는 내용을 다루었다. 특히, 도출된 경로 집단에 순위화 방법론을 적용하여, 신약 개발의 표적으로서 연구할 가치가 있을 것으로 예상되는 물질 및 이를 포함한 경로들을 우선적으로 검토할 수 있게끔 하였다. 이를 위해 단어의 중심성 지표를 활용하는 방법, 관계의 빈도를 이용하는 방식, 관계 빈도의 비율을 적용하는 방안 등 총 세 가지 순위화 기제를 고안하여 활용하였는데, 관계의 빈도를 활용했을 때에 전반부에서 이용되었던 종전의 연구 결과들이 가장 높은 순위에 위치함을 확인할 수 있었다.

    또한, 후반부 연구의 마지막 한 절을 모든 작동 기제가 완전하게 규명되지 않은 WARS1/tryptophan 관련 분야에 본 연구의 방법론을 대입하는 내용에 할당하여, 아직 직접적으로 밝혀지지는 않았지만 지금까지의 연구를 통해 유추만 가능했던 내용을 명시적인 경로로 표현해 보는 새로운 시도를 덧붙였다. 이를 통해, 부분적인 실험과 정황을 통해서 가설로만 제안되었던 내용을 문헌 기반 발견법을 통해 다시 한번 확인함으로써 본 연구 방법론의 유용성을 강화함과 동시에, 신약 개발로 이어질 가능성이 있는 또 다른 경로들을 함께 제시하는 효과를 누릴 수 있었다.

    본 연구는 딥러닝을 활용한 관계 분류 모델이 제한적으로만 활용되었다는 점, 분석 대상 텍스트가 제목 및 초록에 한정되어 있다는 점 등 몇몇 한계점이 존재함에도 불구하고, conceptual biology의 방법론이 신약 표적 물질 발굴에 실질적으로 이용될 수 있음을 여러 근거를 들어 체계적으로 증명했다는 점에서 큰 의의를 지닌다고 볼 수 있다. 본 연구를 기점으로 conceptual biology 방법론이 생물 의학 및 신약 개발 연구에 더욱 보편적으로 활용된다면, 관련 연구 과정상에 발생하는 각종 사회적·경제적 비용을 획기적으로 절감할 수 있을 것으로 기대한다.

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼