The rapid advancement of Generative Artificial Intelligence (Generative AI) has brought innovation across various industries; however, it has also raised growing concerns regarding cultural bias embedded in the language expressions produced by such mo...
The rapid advancement of Generative Artificial Intelligence (Generative AI) has brought innovation across various industries; however, it has also raised growing concerns regarding cultural bias embedded in the language expressions produced by such models. In particular, as Korean-based generative language models are increasingly integrated into real-world applications, numerous cases have been reported where these models generate stereotypical or imbalanced expressions concerning gender, region, social class, or political orientation. This issue goes beyond a mere technical flaw and is regarded as a significant socio-linguistic problem that may distort social perceptions and reproduce discrimination.
Previous studies on bias in language models have primarily focused on English-based large-scale models, while systematic and fine-grained analyses of Korean models remain limited. Moreover, existing evaluation methods have mostly relied on single quantitative metrics such as LPBS or StereoSet, or on binary classifications of bias occurrence, thereby failing to capture the deeper structural relations between socio-cultural context and linguistic expression.
To address these limitations, this study proposes an integrated analytical framework that systematically measures and interprets cultural bias embedded in Korean generative AI language models through both quantitative and qualitative approaches. The framework decomposes bias generation into three structural levels—input, model processing, and output—and analyzes the manifestation of bias within each stage. This multi-layered design enables a deeper understanding of the mechanisms by which subtle forms of cultural representation emerge, beyond what can be detected through numerical metrics alone.
The experiments were conducted using two representative Korean generative language models, KoGPT and KoAlpaca. Five culturally sensitive scenarios relevant to Korean society—gender, regional identity, social hierarchy, political ideology, and family composition—were each represented by six prompts, resulting in a total of 30 prompt sets. Each prompt was generated five times, producing 300 responses per model to ensure response diversity and consistency.
For quantitative analysis, the Language Politeness Bias Score (LPBS) and StereoSet metrics were applied to numerically evaluate bias tendencies and compare the magnitude of bias across scenarios. For qualitative analysis, a panel of three experts specializing in sociolinguistics, AI ethics, and data science conducted manual assessments based on a predefined codebook including criteria such as politeness, stereotypical expression, exclusionary language, and diversity acceptance.
The results revealed meaningful differences in bias patterns and response strategies between the two models. KoGPT tended to produce more polite linguistic forms but exhibited generalized expressions toward specific social groups (e.g., gender, region, or occupation) in approximately 41.7% of its responses. In contrast, KoAlpaca showed a higher proportion of neutral or counter-stereotypical responses (63.5%), but a relatively large number of avoidance or contextually inconsistent answers (18.6%), indicating lower contextual understanding. These findings suggest that surface-level politeness does not necessarily correspond to the absence of cultural bias, underscoring the need for evaluations that jointly consider linguistic style and semantic depth.
Among the tested scenarios, gender, political orientation, and regional identity showed the largest deviations in bias scores, which may be attributed to an imbalance in the cultural diversity of training data. The hybrid integration of quantitative metrics and qualitative coding proved to be crucial in enhancing the validity of bias detection. Notably, several responses judged as neutral in quantitative evaluation were reinterpreted as avoidance-type or latent biases in qualitative coding, highlighting the limitations of single-metric evaluations.
This study contributes both academically and practically by proposing and empirically validating a cultural bias evaluation framework that incorporates sociolinguistic perspectives into the assessment of Korean generative AI models. The multi-layer bias structure analysis, standardized prompt-based scenario design, pseudo-coded evaluation logic, and integrated quantitative–qualitative assessment process provide a foundational methodology for future research on ethical AI development and culturally sensitive model evaluation. Future work will aim to advance this framework toward a more automated and scalable structure, enabling real-time bias detection and fostering generative AI systems that promote inclusive and equitable language generation.