RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    스위스 토너먼트 방식을 활용한 LLM 기반 고난도 질문 자동 선별 방법론 = An LLM-Based Methodology for Automatic Selection of High-Difficulty Questions Using the Swiss Tournament Format

    한글로보기

    https://www.riss.kr/link?id=T17372413

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    거대 언어 모델의 발전에 따라 문맥 내 학습은 언어 모델의 대표적인 활용법으로 주목받으며, 이에 다양한 프롬프트 기법이 연구되고 있다. 특히, 사고 과정의 명시를 유도한 few-shot CoT(Chain-of-Thought)는 소수의 예시 제공만으로 추론 성능을 극대화한 방법으로 알려져 있으나, 예시 구성에 따라 성능 편차가 발생한다는 한계가 존재한다. 기존 연구는 다양성이나 불확실성 등 일부 기준을 중심으로 예시를 구성할 질문을 선정해왔으나, 난이도 기반의 접근은 상대적으로 미진한 실정이다. 이에 본 연구는 언어 모델을 활용한 쌍대 난이도 비교와 스위스 토너먼트 구조를 결합하여 고난도 질문을 체계적으로 선별하고, 이를 인간 주석과 함께 새로운 few-shot CoT 예시를 구축하여 추론을 수행하는 방법론을 제안한다. 제안 방법론 성능 평가를 위해 수학 서술형 벤치마크인 GSM8K 데이터셋 1,319개 문항을 대상으로 실험을 수행한 결과, 제안 방법론이 무작위 선정, 불확실성 기반, 그리고 직접 난이도 평가 방식 대비 각각 2.12%p, 1.36%p, 10.16%p 높은 정확도를 달성하여 우수한 성능을 보임을 확인하였다.
    번역하기

    거대 언어 모델의 발전에 따라 문맥 내 학습은 언어 모델의 대표적인 활용법으로 주목받으며, 이에 다양한 프롬프트 기법이 연구되고 있다. 특히, 사고 과정의 명시를 유도한 few-shot CoT(Chain-o...

    거대 언어 모델의 발전에 따라 문맥 내 학습은 언어 모델의 대표적인 활용법으로 주목받으며, 이에 다양한 프롬프트 기법이 연구되고 있다. 특히, 사고 과정의 명시를 유도한 few-shot CoT(Chain-of-Thought)는 소수의 예시 제공만으로 추론 성능을 극대화한 방법으로 알려져 있으나, 예시 구성에 따라 성능 편차가 발생한다는 한계가 존재한다. 기존 연구는 다양성이나 불확실성 등 일부 기준을 중심으로 예시를 구성할 질문을 선정해왔으나, 난이도 기반의 접근은 상대적으로 미진한 실정이다. 이에 본 연구는 언어 모델을 활용한 쌍대 난이도 비교와 스위스 토너먼트 구조를 결합하여 고난도 질문을 체계적으로 선별하고, 이를 인간 주석과 함께 새로운 few-shot CoT 예시를 구축하여 추론을 수행하는 방법론을 제안한다. 제안 방법론 성능 평가를 위해 수학 서술형 벤치마크인 GSM8K 데이터셋 1,319개 문항을 대상으로 실험을 수행한 결과, 제안 방법론이 무작위 선정, 불확실성 기반, 그리고 직접 난이도 평가 방식 대비 각각 2.12%p, 1.36%p, 10.16%p 높은 정확도를 달성하여 우수한 성능을 보임을 확인하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    With the advancement of large language models (LLMs), in-context learning has emerged as a representative approach, leading to active research on various prompting techniques. In particular, few-shot CoT prompting, which elicits explicit reasoning processes, has demonstrated strong reasoning performance with only a few examples; however, its performance varies depending on the composition of the examples. Previous studies have selected questions for example construction based on criteria such as diversity or uncertainty, while difficulty-based approaches remain relatively underexplored. In this context, this study proposes a methodology that systematically selects high-difficulty questions by combining pairwise difficulty comparisons performed by a language model with a Swiss tournament structure, and constructs new few-shot CoT exemplars with human annotations for reasoning tasks. To evaluate the effectiveness of the proposed methodology, experiments were conducted on 1,319 problems from GSM8K, a mathematical reasoning benchmark. The results confirmed that the proposed method achieved superior performance, with accuracy improvements of 2.12%p, 1.36%p, and 10.16%p over random selection, uncertainty-based selection, and direct difficulty evaluation methods, respectively.
    번역하기

    With the advancement of large language models (LLMs), in-context learning has emerged as a representative approach, leading to active research on various prompting techniques. In particular, few-shot CoT prompting, which elicits explicit reasoning pro...

    With the advancement of large language models (LLMs), in-context learning has emerged as a representative approach, leading to active research on various prompting techniques. In particular, few-shot CoT prompting, which elicits explicit reasoning processes, has demonstrated strong reasoning performance with only a few examples; however, its performance varies depending on the composition of the examples. Previous studies have selected questions for example construction based on criteria such as diversity or uncertainty, while difficulty-based approaches remain relatively underexplored. In this context, this study proposes a methodology that systematically selects high-difficulty questions by combining pairwise difficulty comparisons performed by a language model with a Swiss tournament structure, and constructs new few-shot CoT exemplars with human annotations for reasoning tasks. To evaluate the effectiveness of the proposed methodology, experiments were conducted on 1,319 problems from GSM8K, a mathematical reasoning benchmark. The results confirmed that the proposed method achieved superior performance, with accuracy improvements of 2.12%p, 1.36%p, and 10.16%p over random selection, uncertainty-based selection, and direct difficulty evaluation methods, respectively.

    더보기

    목차 (Table of Contents)

    • 제1장. 서론 1
    • 제2장. 관련 연구 4
    • 2.1. 언어 모델 4
    • 2.2. 문맥 내 학습 5
    • 2.3. Chain-of-Thought 6
    • 제1장. 서론 1
    • 제2장. 관련 연구 4
    • 2.1. 언어 모델 4
    • 2.2. 문맥 내 학습 5
    • 2.3. Chain-of-Thought 6
    • 2.4. Few-shot CoT 예시 선정 7
    • 2.5. 난이도 기반 학습 8
    • 제3장. 제안 방법론 9
    • 3.1. 방법론 개요 9
    • 3.2. 데이터셋 재구성 10
    • 3.3. 난이도 기반 예시 질문 선정 11
    • 3.4. 인간 주석 및 추론 16
    • 제4장. 실험 18
    • 4.1. 실험 개요 18
    • 4.2. 데이터 구축 및 질문 선별 20
    • 4.3. 성능 평가 22
    • 제5장. 결론 25
    • 참고문헌 27
    • Abstract 36
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼