The purpose of this study is to establish the quantitative methodology for Korean co-occurrence research. This thesis shows the quantitative way for the study of a collocation and a colligation in Korean. For this study, this thesis uses the Sejong pa...
The purpose of this study is to establish the quantitative methodology for Korean co-occurrence research. This thesis shows the quantitative way for the study of a collocation and a colligation in Korean. For this study, this thesis uses the Sejong part-of-speech tagged corpus on a 5,500,00 word scale and the Korea-1 syntactic-relation tagged corpus on a 10 million word scale.
This thesis is divided up into seven chapters. The chapter Ⅰ is an introduction to this study, summarizes the antecedent study and make a statement of the object of our study.
In chapter Ⅱ, the quantitative way of the co-occurrence research is established and corpora which are used in this study are described. Firstly, for the exact and objective quantitative study, antecedent statistic formulas are examined with Korean data. And we attempted new statistic formula which is base on Possion distribution. After these examination, we choose some formulas which are most proper for significant co-occurrence measuring. Secondly, for measuring the restrict co-occurrence, we apply the mean(of position score) and the sample deviation(of positon score) to consideration of restricted co-occurrence position. Thirdly, the factor analysis is described for co-occurrence patterns among collocations. And then, the corpora processes are described.
The chapter Ⅲ is described for the characteristic of adjacent co-occurrence of base word. The base words for this study are selected from the POS tagged corpus on a 5.5 million word scale. Every selected word is high frequent and representable in each POS. With these words, the adjacent co-occurrence is analyzed in the view of the collocation and the colligation. The result of this study shows the different linguistic character of words in same POS.
The chapter Ⅳ is described for the characteristic of syntactic co-occurrence of base word. The base words for this study are selected from the syntactic tagged corpus on a 10 million word scale. This corpus is tagged the modify relation for the noun and verb. And for the verb, the corpus is tagged the argument relation. With this corpus, we can investigate the argument occurrence patterns.
The chapter Ⅴ is described for the characteristic of co-occurrence patterns among collocations. The base words for this study are selected from the POS tagged corpus on a 5.5 million word scale. But differently to Ch. Ⅲ, this chapter uses 299 separated files for the proper factor analysis. This approach can be useful tool for semi-automatic identification of underlying word senses and uses.
In chapter Ⅵ, we compare the quantitative ways which are used in Ch. Ⅲ, Ch. Ⅳ, and Ch. Ⅴ with same word '말하다(malha-da: 'say')'. With this comparing, we can arrange the differences of each quantitative approach and show the strong point of each quantitative approach explicitly.
The chapter Ⅶ describe the summary of this thesis and explicate the remaining Issues.
The Quantitative study on co-occurrence in Korean is meaningful that new methodology in Korean quantitative research are shown in collocations and colligations. This kind of studies haven't shown in Korean study before this thesis. So this study can indicate the new approach way for the theoretical Korean study, the natural language processing, and the teaching second language etc.