This paper reports the author identification (or authorship attribution) study of Korean Corpus A and Corpus B by the quantitative analysis of linguistic characteristics, consisting of letters and symbols, phrases, morpheme tags, and n-grams derived f...
This paper reports the author identification (or authorship attribution) study of Korean Corpus A and Corpus B by the quantitative analysis of linguistic characteristics, consisting of letters and symbols, phrases, morpheme tags, and n-grams derived from morphemes. Minimum distance methods(Euclidean, Chisquare, Weighted Euclidean, Cosine, Symmetric Kullback-Leibler) and machine learning methods such as k-nearest neighbor(k-NN), support vector machine (SVM), random forests(RF) are applied to linguistic features extracted from the classified texts. Results show that SVM and RF are superior in recall and precision compared to k-NN and minimum distance methods with the discriminant rate as high as 98% in Corpus A and 100% in Corpus B. Among five distances considered, Weighted Euclidean and Symmetric Kullback-Leibler distances are better than the others.