The purpose of this study is to build a construction dataset that considers the syntactic and semantic characteristics of language and to use it in the field of Natural Language Processing(NLP). Currently, datasets used in the NLP domain are mainly f...
The purpose of this study is to build a construction dataset that considers the syntactic and semantic characteristics of language and to use it in the field of Natural Language Processing(NLP). Currently, datasets used in the NLP domain are mainly focused on building large amounts of corpus, and language models that have learned such datasets do not properly handle constructions. This thesis attempted to describe the process of building a dataset based on construction grammar and to reveal the effectiveness of the dataset built through experiments using an actual NLP language model, BERT.
The subject of this study is the construction of compound directional complement in Chinese. The meanings of this construction are rich and various, but some of the constructions appears only in the corpus of certain fields, making it difficult for language models to learn sufficiently with the general corpus given to them. In this thesis, construction data was extracted from various areas of BCC corpus, CCL corpus, and colloquial corpus, and processed into a form that language models can learn to build a compound directional complement construction dataset of about 250,000 sentences.
Chapter 2 summarizes the information of the construction that the dataset must contain before constructing the dataset. In particular, the syntactic and semantic characteristics of each compound directional complement construction were analyzed. The compound directional complements defined in this thesis appear in six forms and has four semantic characteristics, and each phrase performs various aspectual functions, such as adding directional colors to the sentence or expressing completion, result, duration, beginning, etc.
Chapter 3 proposes a method of search and extraction based on the forms of the construction, and describes the process of constructing an actual construction dataset from three source corpora. The processes of constructing dataset were divided into first and second extraction process for each corpus so that the constructions with various forms and meanings can be included. In the first section of extraction, the search formulas were divided into four types to describe the process of extracting examples considering various forms of construction. In the second section of extraction, only actual construction data was selected and processed among the extracted sentences using a part-of-speech(POS) tagging program.
Chapter 4 deals with quantified data, compositional information, and distributional characteristics of construction dataset actually constructed through extraction methods. In addition, the chapter describes the results of statistical analysis such as the ratio of complement and type among dataset and the analysis result of the information of frequently combined verb of each construction. As a result of the analysis, it was verified that various semantic and syntactic information of the construction was appropriately included in the construction dataset.
In Chapter 5, experiments using two types of BERT models were conducted to confirm the effectiveness of the actual dataset. As a result of learning and testing the dataset built for the model, it was confirmed that accuracy, precision, recall, and F1 score all increased significantly compared to the pre-trained model. Furthermore, the predictions of each model through some actual test examples were examined and it was confirmed that the constructed dataset contains sufficient information of the construction, and that learning the construction datasets can help improve the processing performance of the NLP models.