거대 언어 모델의 발전에 따라 문맥 내 학습은 언어 모델의 대표적인 활용법으로 주목받으며, 이에 다양한 프롬프트 기법이 연구되고 있다. 특히, 사고 과정의 명시를 유도한 few-shot CoT(Chain-o...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17372413
서울 : 국민대학교 비즈니스IT전문대학원, 2025
학위논문(석사) -- 국민대학교 비즈니스IT전문대학원 , 비즈니스IT전공 , 2026. 2
2025
한국어
서울
ⅳ, 36 ; 26 cm
지도교수: 김남규
I804:11014-200000960552
0
상세조회0
다운로드거대 언어 모델의 발전에 따라 문맥 내 학습은 언어 모델의 대표적인 활용법으로 주목받으며, 이에 다양한 프롬프트 기법이 연구되고 있다. 특히, 사고 과정의 명시를 유도한 few-shot CoT(Chain-o...
거대 언어 모델의 발전에 따라 문맥 내 학습은 언어 모델의 대표적인 활용법으로 주목받으며, 이에 다양한 프롬프트 기법이 연구되고 있다. 특히, 사고 과정의 명시를 유도한 few-shot CoT(Chain-of-Thought)는 소수의 예시 제공만으로 추론 성능을 극대화한 방법으로 알려져 있으나, 예시 구성에 따라 성능 편차가 발생한다는 한계가 존재한다. 기존 연구는 다양성이나 불확실성 등 일부 기준을 중심으로 예시를 구성할 질문을 선정해왔으나, 난이도 기반의 접근은 상대적으로 미진한 실정이다. 이에 본 연구는 언어 모델을 활용한 쌍대 난이도 비교와 스위스 토너먼트 구조를 결합하여 고난도 질문을 체계적으로 선별하고, 이를 인간 주석과 함께 새로운 few-shot CoT 예시를 구축하여 추론을 수행하는 방법론을 제안한다. 제안 방법론 성능 평가를 위해 수학 서술형 벤치마크인 GSM8K 데이터셋 1,319개 문항을 대상으로 실험을 수행한 결과, 제안 방법론이 무작위 선정, 불확실성 기반, 그리고 직접 난이도 평가 방식 대비 각각 2.12%p, 1.36%p, 10.16%p 높은 정확도를 달성하여 우수한 성능을 보임을 확인하였다.
다국어 초록 (Multilingual Abstract)
With the advancement of large language models (LLMs), in-context learning has emerged as a representative approach, leading to active research on various prompting techniques. In particular, few-shot CoT prompting, which elicits explicit reasoning pro...
With the advancement of large language models (LLMs), in-context learning has emerged as a representative approach, leading to active research on various prompting techniques. In particular, few-shot CoT prompting, which elicits explicit reasoning processes, has demonstrated strong reasoning performance with only a few examples; however, its performance varies depending on the composition of the examples. Previous studies have selected questions for example construction based on criteria such as diversity or uncertainty, while difficulty-based approaches remain relatively underexplored. In this context, this study proposes a methodology that systematically selects high-difficulty questions by combining pairwise difficulty comparisons performed by a language model with a Swiss tournament structure, and constructs new few-shot CoT exemplars with human annotations for reasoning tasks. To evaluate the effectiveness of the proposed methodology, experiments were conducted on 1,319 problems from GSM8K, a mathematical reasoning benchmark. The results confirmed that the proposed method achieved superior performance, with accuracy improvements of 2.12%p, 1.36%p, and 10.16%p over random selection, uncertainty-based selection, and direct difficulty evaluation methods, respectively.
목차 (Table of Contents)