Amid the decline in the school-age population, universities are confronting a shift in the composition of their educational clientele. In response, efforts to recruit international students have intensified, accompanied by the diversification of instr...
Amid the decline in the school-age population, universities are confronting a shift in the composition of their educational clientele. In response, efforts to recruit international students have intensified, accompanied by the diversification of instructional environments. In classroom settings composed primarily of international students with limited Korean communicative competence, the use of real-time speech recognition–based speech-to-text (STT) translation systems has become virtually indispensable. This study analyzes such systems as employed in actual classroom contexts, examining translation patterns and errors in order to identify their current limitations.
The findings reveal that real-time STT translation systems exhibit errors across lexical, phonological, and grammatical categories. Within the lexical category, error types include contextually inappropriate word choices, homonyms, polysemy, technical terminology, idiomatic expressions, proper nouns, unnecessary translations, mixed usage of English and numerals, loanword interference, and abbreviations.
In the phonological category, errors arise from the recognition of consonants and vowels, the pronunciation of loanwords, and pauses occurring during speech production. While such errors are typically absent in text-based machine translation, they emerge in speech recognition translation due to the real-time processing of spontaneous spoken language.
Grammatical errors occur less frequently than lexical or phonological ones; however, prominent issues include the failure to appropriately resolve omitted subjects in Korean—particularly in selecting corresponding pronouns—as well as errors involving conjunctions and word order.
These diverse error types can be attributed, in essence, to a lack of contextual understanding. In written texts, context is relatively constrained, consisting of the information provided within the discourse. In contrast, real-time spoken interaction involves not only linguistic content but also situational context. Humans can integrate multimodal cues—such as auditory signals, facial expressions, gestures, and nuanced intonation—to interpret meaning, whereas machines rely solely on probabilistic computation and the analysis of input data, limiting their capacity to process such complex contextual information. Therefore, continued advancements in speech recognition technology remain necessary.
It is anticipated that the linguistic error phenomena identified in this study will provide practical guidelines for instructors seeking to utilize such systems effectively, and further contribute to the ongoing improvement and development of speech recognition translation technologies.