Despite many studies showing that negation affects the performance of large language models (LLMs), there exist very few benchmarks that explicitly evaluate models’ negation understanding. In particular, benchmarks targeting negation comprehension i...
Despite many studies showing that negation affects the performance of large language models (LLMs), there exist very few benchmarks that explicitly evaluate models’ negation understanding. In particular, benchmarks targeting negation comprehension in Korean remain scarce.
In this work, we first conduct a detailed analysis of Korean negation phenomena and sentence structures. We then examine KoNUBench, a novel benchmark designed to assess LLMs’ ability to understand negation in Korean. KoNUBench is a multiple-choice dataset in which models must select the correct standard negation from various distractor options—including local negation, contradiction, and paraphrase—thereby requiring sentence-level understanding of negation.
We further validate KoNUBench as a reliable benchmark for measuring LLMs’ negation understanding in Korean, and analyze the performance of 43 LLMs to provide a comprehensive evaluation of their negation reasoning capabilities.