Automatic dialect identification is a technology that automatically distinguishes dialectal characteristics of speakers from speech signals, representing an interdisciplinary research field that should be developed upon the theoretical foundation of d...
Automatic dialect identification is a technology that automatically distinguishes dialectal characteristics of speakers from speech signals, representing an interdisciplinary research field that should be developed upon the theoretical foundation of dialectology. However, existing studies have focused solely on improving the performance of learning models without connecting to dialectological theories. This study aims to analyze the linguistic characteristics of Korean dialect speech based on the framework of Korean dialectology and systematically apply these findings to automatic dialect identification systems.
To conduct this research, we first redefined dialect classes by applying the dialectal region classification system based on Korean dialectological criteria, departing from the existing administrative district-based dialect classification. To identify phonetic differences among dialects, vowel analysis and eGeMAPS feature analysis were performed. In vowel analysis, vowel formants (F1, F2) and duration were extracted from each dialectal region's speech, and dialectal differences in vowels and phonological change phenomena were quantitatively analyzed through analysis of variance (ANOVA). In eGeMAPS feature analysis, acoustic features in frequency, energy, spectral, and temporal domains were extracted for pairwise contrastive analysis between dialectal regions. The results confirmed that frequency and energy-related features contribute statistically significantly to dialect discrimination.
The automatic dialect identification system was implemented using two approaches. In the feature-based approach, binary classification models were constructed using eGeMAPS acoustic features that showed significant differences between dialects, and SHAP (SHapley Additive exPlanations) analysis was applied to interpret the relative importance and linguistic significance of acoustic features contributing to dialect prediction. In the self-supervised learning approach, end-to-end dialect identification was performed using the XLS-R (Cross-lingual Speech Representations) model. Additionally, a novel concept of "dialect density" was introduced to visualize speech segments that the model focuses on during dialect identification, and the concordance with dialectological tagging results was measured to enhance model interpretability.
This study holds academic significance in systematically integrating dialectological theories into automatic dialect identification research. This interdisciplinary approach leads to contributions in both academic and practical domains. Academically, it presents a new research direction that can objectively validate the validity of traditional phonology-based dialect demarcation through computational modeling methodologies. Practically, it establishes a foundation for tools that can provide comprehensible evidence for investigators in forensic science when estimating the regional origin of unidentified speakers. The results of this study are expected to contribute to the quantification of Korean dialectology research and the strengthening of the theoretical foundation for automatic dialect identification technology.