Conceptual biology is a new research trend in the field of biomedical science that seeks to gain
new insights by collecting fragmented knowledge by reviewing the findings of biomedical fields
accumulated in multiple databases and linking related conce...
Conceptual biology is a new research trend in the field of biomedical science that seeks to gain
new insights by collecting fragmented knowledge by reviewing the findings of biomedical fields
accumulated in multiple databases and linking related concepts. The concept of conceptual biology
has also been applied to modern drug developments, with the rapid development of text mining
techniques to extract desired information from biomedical data, such as paper publications,
systematically well managed.
However, conceptual biology is still not the main methodology for biomedical and new drug
development research. This is because there is a deep-rooted perception among researchers that
hypotheses generated and proposed in an automated manner through biomedical text mining
technology are less reliable than those made by conventional research methods. Therefore, prior to
recklessly introducing new hypotheses and drug targets under the banner of conceptual biology, it
is necessary to show that the research results obtained through the existing experimental-oriented
research methods can be derived equally by computerized text mining techniques.
Therefore, in this work, we first demonstrate how reliable the connections between entities
extracted from unstructured text data are through comparisons with existing research cases, and
then apply the methodology to actual new drug target discovery to propose new hypotheses and
materials to researchers. In this process, we used the theoretical framework of literature-based
discovery (LBD) to bridge between conceptual biology and the field of library and information
science, and used an enzyme called aminoacyl-tRNA synthetase (ARS), which has recently been
spotlighted as a criterion for finding drug target proteins, as a key research material.
To implement the above intentions and configurations in practice, the research procedure was
also conducted in two major steps. First, in the first part, we summarized the existing research
results related to ARS and expressed them in the form of a series of paths, and then, examined
whether they could be reproduced in the same way through the currently used text mining
technique. In this process, research papers related to the three pairs of ARS and amino acids
(LARS1/leucine, QARS1/glutamine, MARS1/methionine) were used as the standard. And by
comparing the results when the literature related only to the contents of LARS1/leucine,
QARS1/glutamine, and MARS1/methionine were used separately with when all the literatures
were combined, we wanted to visually identify the benefits of increasing the size of the literature
group to be analyzed.
As a result of the experiments, it was confirmed that most of the contents of the standard papers
were reproduced through text mining techniques, demonstrating that this research method is fully
utilizable unlike researchers' stereotypes. In particular, it is shown that the papers are better
represented when using the entire data compared to utilizing only the literature groups directly
related to the hypothesis we want to generate, indicating that even if we aim to generate
hypotheses in a particular ARS field, it is better to utilize the data around it. Also, throughout this
process, we also coordinate and optimize the specifications of the methodology, including the type
of named entity dictionary to be used, the stop words list to be included, and the categories of data
to be analyzed, to lay the foundation for what will follow.
Meanwhile, in the latter part of the study, we explored substances that mediated the interaction
between ARS and amino acids, and proposed new features of ARS which have not yet been
proven. Especially, ranking algorithm was applied to the derived paths, allowing researchers to
preferentially review materials that are expected to be worth studying as targets for new drug
development. To this end, a total of three ranking mechanisms were devised and utilized: using the
centrality indicators of words, considering the frequency of relations, and applying the ratio of
relation frequency. And, as a result, in case of the second method, utilizing the frequency of
relations, the results of previous studies in the first half were found to be at the highest rank.
In addition, the last one section of the latter part of the study was assigned to the contents of
substituting our methodology to the field of WARS1/tryptophan, where all operating mechanisms
were not fully elucidated, being added a new attempt to express in an explicit path what has not
yet been directly identified. Through this, we reinforced the usefulness of our research
methodology by reconfirming the contents suggested only as hypotheses through partial
experiments and contexts. At the same time, we were also able to enjoy the effect of presenting
other possible paths that could lead to the development of new drugs.
Despite the existence of some limitations, such as limited use of relation classification models
using deep learning and incompleteness of text subject to analysis solely on titles and abstracts,
this work is significant in that it systematically proves that the methodology of conceptual biology
can be used in drug target discovery on several grounds. Starting with this study, if conceptual
biology is more commonly used for biomedical and drug development research, it is expected that
various social and economic costs incurred in the related research process will be dramatically
reduced.