This study evaluates the safety of large language models (LLMs) not as a single metric, but as a multidimensional set of safety characteristics. Prior evaluations have primarily emphasized isolated indicators, such as refusal of harmful requests or su...
This study evaluates the safety of large language models (LLMs) not as a single metric, but as a multidimensional set of safety characteristics. Prior evaluations have primarily emphasized isolated indicators, such as refusal of harmful requests or suppression of toxic language. However, such approaches do not adequately capture the complexity of safety judgments required in real-world service settings. In practice, LLM safety depends on the interplay of factual reliability, bias mitigation, robustness to jailbreak and policy-circumvention attempts, the ability to maintain normative constraints under sustained user pressure, and context-sensitive judgment that accounts for age-related considerations. These challenges are particularly salient in Korean-language environments, where linguistic and cultural factors are insufficiently reflected in evaluation frameworks developed primarily for English-language settings. Accordingly, this study conducts a controlled comparative evaluation of representative Korean and global LLMs, empirically identifies where model-specific safety strengths and vulnerabilities emerge across evaluation dimensions, and suggests directions for safety assessment better aligned with Korean deployment contexts.
To this end, we evaluate two Korean-developed models—NAVER’s HyperCLOVA-family model hcx-003 and Kakao’s KanaNA-family model kanana-nano-2.1b—and two global models—OpenAI’s GPT-4o-family model gpt-4o-mini and Anthropic’s Claude 4-family model claude-sonnet-4—using a comprehensive suite of benchmarks. The evaluation covers truthfulness (TruthfulQA), jailbreak vulnerability (AdvBench), susceptibility to user pressure and sycophancy (SYCON Bench), Korean-language bias and ethical reasoning (SQuARe, KoSBi, KoBBQ), and an age-rating classification dataset constructed in this study.
The results show that none of the four models consistently outperforms the others across safety metrics, and that safety profiles differ substantially by model, language, and evaluation dimension. Notably, across multiple benchmarks, even the same model exhibits meaningfully different safety behaviors in English versus Korean settings.
On global benchmarks (TruthfulQA, AdvBench, and SYCON Bench), global models generally outperform Korean-developed models. claude-sonnet-4 and gpt-4o-mini achieve high truthfulness and informativeness on TruthfulQA. In particular, claude-sonnet-4 records the highest Turn of Flip (ToF) on SYCON Bench, indicating the strongest ability to maintain a normative stance under sustained user pressure. This tendency is more pronounced in English prompts, where most models show higher ToF values than in Korean, suggesting that models respond with more explicitly stated and stance-consistent positions in English contexts.
In contrast, ToF scores decrease in Korean settings even for global models, which may reflect an adaptive shift toward more user-accommodating conversational behavior in Korean dialogue contexts. Taken together, these findings suggest that sycophancy is not a fixed model attribute, but a safety-relevant behavior that can vary dynamically with linguistic and cultural context.
Korean-specific benchmarks (SQuARe, KoSBi, and KoBBQ) further clarify these differences. While global models maintain relatively stable normative compliance in Korean, performance gaps narrow compared to English benchmarks. Moreover, KoBBQ and KoSBi indicate that vulnerabilities are shaped less by a simple Korean-versus-global performance hierarchy than by the type of bias being tested and the surrounding social context. This highlights the limitations of relying solely on English-centric benchmarks to assess LLM safety in Korean-language environments.
Age-rating classification most clearly reveals systematic differences in safety behavior. claude-sonnet-4 separates the 12-, 15-, and 19-year categories with relatively stable boundaries, suggesting that norm-based judgment generalizes consistently to age-contextual settings. gpt-4o-mini likewise preserves continuity across adjacent age categories and produces error patterns that remain interpretable from a policy perspective. In contrast, Korean-developed models exhibit pronounced limitations. HyperCLOVA (hcx-003) maintains reasonable discrimination for low-risk (ages 0 and 7) and high-risk (age 19) content, but shows concentrated errors in the mid-tier categories (ages 12 and 15), reflecting a tendency to underestimate risk. KanaNA (kanana-nano-2.1b) achieves extremely low jailbreak attack success rates, indicating a conservative safety posture; however, its performance degrades on age-rating classification, suggesting that stronger filtering may simultaneously reduce the model’s capacity for nuanced reasoning and norm-consistent judgment.
In contrast, domestic models exhibited pronounced structural limitations. HyperCLOVA (hcx-003) showed reasonable discrimination for low-risk (ages 0 and 7) and clearly high-risk (age 19) content, but concentrated errors in the intermediate categories (ages 12 and 15), revealing a systematic under-rating bias. KanaNA (kanana-nano-2.1b), while achieving extremely low attack success rates on jailbreak benchmarks, suffered substantial degradation in age-rating classification, informativeness, and bias-related metrics. This suggests that aggressive pre-filtering can suppress not only harmful outputs but also the model’s capacity for nuanced reasoning and normative judgment.
In summary, this study provides empirical evidence that LLM safety cannot be reduced to overall model capability, but should be understood as a multidimensional property shaped by language environment, evaluation criteria, and policy context. Truthfulness, normative consistency, adversarial robustness, and age-contextual judgment represent complementary dimensions of safety, and improvements in one area may not fully compensate for vulnerabilities in others. These findings support the need for multidimensional, multi-benchmark evaluation frameworks and motivate layered safety management strategies that integrate evaluation outcomes with operational policy requirements.