RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    KoNUBench: Measuring LLM's Negation Understanding in Korean = KoNUBench: 대규모 언어모델의 한국어 부정 표현 이해 능력 평가를 위한 벤치마크

    한글로보기

    https://www.riss.kr/link?id=T17450539

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Despite many studies showing that negation affects the performance of large language models (LLMs), there exist very few benchmarks that explicitly evaluate models’ negation understanding. In particular, benchmarks targeting negation comprehension in Korean remain scarce.
    In this work, we first conduct a detailed analysis of Korean negation phenomena and sentence structures. We then examine KoNUBench, a novel benchmark designed to assess LLMs’ ability to understand negation in Korean. KoNUBench is a multiple-choice dataset in which models must select the correct standard negation from various distractor options—including local negation, contradiction, and paraphrase—thereby requiring sentence-level understanding of negation.
    We further validate KoNUBench as a reliable benchmark for measuring LLMs’ negation understanding in Korean, and analyze the performance of 43 LLMs to provide a comprehensive evaluation of their negation reasoning capabilities.
    번역하기

    Despite many studies showing that negation affects the performance of large language models (LLMs), there exist very few benchmarks that explicitly evaluate models’ negation understanding. In particular, benchmarks targeting negation comprehension i...

    Despite many studies showing that negation affects the performance of large language models (LLMs), there exist very few benchmarks that explicitly evaluate models’ negation understanding. In particular, benchmarks targeting negation comprehension in Korean remain scarce.
    In this work, we first conduct a detailed analysis of Korean negation phenomena and sentence structures. We then examine KoNUBench, a novel benchmark designed to assess LLMs’ ability to understand negation in Korean. KoNUBench is a multiple-choice dataset in which models must select the correct standard negation from various distractor options—including local negation, contradiction, and paraphrase—thereby requiring sentence-level understanding of negation.
    We further validate KoNUBench as a reliable benchmark for measuring LLMs’ negation understanding in Korean, and analyze the performance of 43 LLMs to provide a comprehensive evaluation of their negation reasoning capabilities.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    대규모 언어 모델(LLM)의 성능에 부정 표현이 영향을 미친다는 연구는 다수 존재하지만, 모델의 부정 이해 능력을 명시적으로 평가하는 벤치마크는 매우 제한적이다. 특히 한국어에서의 부정 이해를 측정하는 벤치마크는 거의 존재하지 않는다.
    본 연구에서는 먼저 한국어의 부정 표현과 문장 구조를 체계적으로 분석한다. 이어서 한국어에서 LLM의 부정 이해 능력을 평가하기 위해 설계된 새로운 벤치마크인 KoNUBench를 분석한다. KoNUBench는 표준 부정을 정답으로 선택하도록 구성된 다지선다형 데이터셋으로, 국소 부정(local negation), 모순(contradiction), 패러프레이즈(paraphrase)와 같은 다양한 오답 선택지들로부터 정답을 구분해야 하므로 문장 수준의 부정 이해를 요구한다.
    또한 KoNUBench가 한국어에서 LLM의 부정 이해 능력을 측정하기 위한 신뢰할 수 있는 벤치마크로 기능함을 검증하고, 43개의 LLM을 대상으로 실험을 수행하여 모델들의 부정 추론 능력을 종합적으로 분석한다.
    번역하기

    대규모 언어 모델(LLM)의 성능에 부정 표현이 영향을 미친다는 연구는 다수 존재하지만, 모델의 부정 이해 능력을 명시적으로 평가하는 벤치마크는 매우 제한적이다. 특히 한국어에서의 부정...

    대규모 언어 모델(LLM)의 성능에 부정 표현이 영향을 미친다는 연구는 다수 존재하지만, 모델의 부정 이해 능력을 명시적으로 평가하는 벤치마크는 매우 제한적이다. 특히 한국어에서의 부정 이해를 측정하는 벤치마크는 거의 존재하지 않는다.
    본 연구에서는 먼저 한국어의 부정 표현과 문장 구조를 체계적으로 분석한다. 이어서 한국어에서 LLM의 부정 이해 능력을 평가하기 위해 설계된 새로운 벤치마크인 KoNUBench를 분석한다. KoNUBench는 표준 부정을 정답으로 선택하도록 구성된 다지선다형 데이터셋으로, 국소 부정(local negation), 모순(contradiction), 패러프레이즈(paraphrase)와 같은 다양한 오답 선택지들로부터 정답을 구분해야 하므로 문장 수준의 부정 이해를 요구한다.
    또한 KoNUBench가 한국어에서 LLM의 부정 이해 능력을 측정하기 위한 신뢰할 수 있는 벤치마크로 기능함을 검증하고, 43개의 LLM을 대상으로 실험을 수행하여 모델들의 부정 추론 능력을 종합적으로 분석한다.

    더보기

    목차 (Table of Contents)

    • 1. Introduction 1
    • 2. Related Work 3
    • 2.1. Challenges of LLMs in Negation Handling 3
    • 2.2. Datasets and Benchmarks for Negation 4
    • 3. Syntactic Negation in Korean 6
    • 1. Introduction 1
    • 2. Related Work 3
    • 2.1. Challenges of LLMs in Negation Handling 3
    • 2.2. Datasets and Benchmarks for Negation 4
    • 3. Syntactic Negation in Korean 6
    • 3.1. Types of Syntactic Negation in Korean 6
    • 3.1.1. An Type(안 계열) 7
    • 3.1.2. Mot Type(못 계열) 7
    • 3.1.3. Malda Type(말다) 8
    • 4. Sentence Structure in Korean 9
    • 4.1. Simple Sentences 9
    • 4.2. Complex Sentences 9
    • 4.2.1. Connected Sentences 10
    • 4.2.2. Sentences with Embedded Clause 10
    • 5. KoNUBench 13
    • 5.1. Task Overview 13
    • 5.2. Capturing the Linguistic Properties of Korean 14
    • 5.2.1. Typology and Distribution of negation in KoNUBench 14
    • 5.2.2 Typology of local negation 15
    • 6. Experiments 18
    • 6.1. Experimental Setup 18
    • 6.1.1. Models 18
    • 6.1.2. Evaluation Method 18
    • 6.1.3. Evaluation and Supervised-fine-tuning 19
    • 6.2. Results 19
    • 6.2.1. Model Size 20
    • 6.2.2. Base Models vs. Instruct-tuned models 20
    • 6.2.3. Effect of Fine-Tuning on Negation Understanding 24
    • 7. Conclusion 25
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼