RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    UAPRM: Uncertainty-Aware Process Reward Models = UAPRM: 불확실성 인지 기반 프로세스 보상 모델

    한글로보기

    https://www.riss.kr/link?id=T17451124

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Process Reward Models (PRMs) have become essential for enhancing the reasoning capabilities of Large Language Models (LLMs) by providing fine-grained feedback at each intermediate step. Recently, generative verifiers have emerged as a promising direction, leveraging the inherent reasoning power of LLMs to verify solutions via Chain-of-Thought (CoT). However, existing generative PRMs suffer from overconfidence, often assigning extreme probability scores to reasoning steps, which undermines their reliability in downstream applications. In this thesis, we introduce the Uncertainty-Aware Process Reward Model (UAPRM), a novel framework designed to mitigate this issue by explicitly modeling confidence within the verification process. We propose a methodology to train generative verifiers that output both a verification rationale and a calibrated confidence score, utilizing a self-consistency-based data augmentation pipeline. Furthermore, to address the inherent data imbalance where high-confidence samples dominate, we employ a token-wise weighted loss strategy that forces the model to learn representations of uncertainty more effectively. Our experiments on mathematical reasoning benchmarks demonstrate that UAPRM outperforms baseline generative verifiers in terms of verification precision and calibration error. We show that these calibrated confidence scores translate into superior performance in Best-of-N selection tasks, establishing UAPRM as a more reliable and robust verifier for complex reasoning.
    번역하기

    Process Reward Models (PRMs) have become essential for enhancing the reasoning capabilities of Large Language Models (LLMs) by providing fine-grained feedback at each intermediate step. Recently, generative verifiers have emerged as a promising direct...

    Process Reward Models (PRMs) have become essential for enhancing the reasoning capabilities of Large Language Models (LLMs) by providing fine-grained feedback at each intermediate step. Recently, generative verifiers have emerged as a promising direction, leveraging the inherent reasoning power of LLMs to verify solutions via Chain-of-Thought (CoT). However, existing generative PRMs suffer from overconfidence, often assigning extreme probability scores to reasoning steps, which undermines their reliability in downstream applications. In this thesis, we introduce the Uncertainty-Aware Process Reward Model (UAPRM), a novel framework designed to mitigate this issue by explicitly modeling confidence within the verification process. We propose a methodology to train generative verifiers that output both a verification rationale and a calibrated confidence score, utilizing a self-consistency-based data augmentation pipeline. Furthermore, to address the inherent data imbalance where high-confidence samples dominate, we employ a token-wise weighted loss strategy that forces the model to learn representations of uncertainty more effectively. Our experiments on mathematical reasoning benchmarks demonstrate that UAPRM outperforms baseline generative verifiers in terms of verification precision and calibration error. We show that these calibrated confidence scores translate into superior performance in Best-of-N selection tasks, establishing UAPRM as a more reliable and robust verifier for complex reasoning.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    프로세스 보상 모델(Process Reward Models, PRMs)은 각 중간 단계에서 세밀한 피드백을 제공함으로써 거대 언어 모델(Large Language Models, LLMs)의 추론 능력을 향상시키는 필수적인 요소가 되었다. 최근에는 생각의 사슬(Chain-of-Thought, CoT)을 통해 LLM 고유의 추론 능력을 활용하여 해답을 검증하는 생성형 검증기가 유망한 연구 방향으로 부상하고 있다. 그러나 기존의 생성형 PRM은 과신 문제를 겪고 있으며, 종종 추론 단계에 극단적인 확률 점수를 부여하여 후속 응용에서의 신뢰성을 저해한다. 본 논문에서는 검증 과정 내에서 신뢰도를 명시적으로 모델링하여 이 문제를 완화하는 새로운 프레임워크인 불확실성 인지 기반 프로세스 보상 모델(UAPRM)을 제안한다. 본 연구에서는 자기 일관성 기반의 데이터 증강 파이프라인을 활용하여, 검증 근거와 보정된 신뢰도 점수를 모두 출력하는 생성형 검증기 학습 방법론을 제시한다. 또한, 고신뢰도 샘플이 지배적인 내재적 데이터 불균형 문제를 해결하기 위해, 모델이 불확실성 표현을 보다 효과적으로 학습하도록 유도하는 토큰 단위 가중 손실 전략을 적용한다. 수학적 추론 벤치마크에 대한 실험을 통해, UAPRM이 검증 정밀도 및 보정 오차 측면에서 기존 생성형 검증기보다 우수한 성능을 보임을 입증하였다. 나아가 이러한 보정된 신뢰도 점수가 Best-of-N 선택 작업에서 뛰어난 성능 향상으로 이어짐을 보임으로써, UAPRM이 복잡한 추론을 위한 보다 신뢰할 수 있고 견고한 검증기임을 확인하였다.
    번역하기

    프로세스 보상 모델(Process Reward Models, PRMs)은 각 중간 단계에서 세밀한 피드백을 제공함으로써 거대 언어 모델(Large Language Models, LLMs)의 추론 능력을 향상시키는 필수적인 요소가 되었다. 최근...

    프로세스 보상 모델(Process Reward Models, PRMs)은 각 중간 단계에서 세밀한 피드백을 제공함으로써 거대 언어 모델(Large Language Models, LLMs)의 추론 능력을 향상시키는 필수적인 요소가 되었다. 최근에는 생각의 사슬(Chain-of-Thought, CoT)을 통해 LLM 고유의 추론 능력을 활용하여 해답을 검증하는 생성형 검증기가 유망한 연구 방향으로 부상하고 있다. 그러나 기존의 생성형 PRM은 과신 문제를 겪고 있으며, 종종 추론 단계에 극단적인 확률 점수를 부여하여 후속 응용에서의 신뢰성을 저해한다. 본 논문에서는 검증 과정 내에서 신뢰도를 명시적으로 모델링하여 이 문제를 완화하는 새로운 프레임워크인 불확실성 인지 기반 프로세스 보상 모델(UAPRM)을 제안한다. 본 연구에서는 자기 일관성 기반의 데이터 증강 파이프라인을 활용하여, 검증 근거와 보정된 신뢰도 점수를 모두 출력하는 생성형 검증기 학습 방법론을 제시한다. 또한, 고신뢰도 샘플이 지배적인 내재적 데이터 불균형 문제를 해결하기 위해, 모델이 불확실성 표현을 보다 효과적으로 학습하도록 유도하는 토큰 단위 가중 손실 전략을 적용한다. 수학적 추론 벤치마크에 대한 실험을 통해, UAPRM이 검증 정밀도 및 보정 오차 측면에서 기존 생성형 검증기보다 우수한 성능을 보임을 입증하였다. 나아가 이러한 보정된 신뢰도 점수가 Best-of-N 선택 작업에서 뛰어난 성능 향상으로 이어짐을 보임으로써, UAPRM이 복잡한 추론을 위한 보다 신뢰할 수 있고 견고한 검증기임을 확인하였다.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 2 Background and Related Work 4
    • 2.1 Process Reward Models 4
    • 2.2 Uncertainty Quantification in LLMs 5
    • 2.3 Inference-Time Scaling 6
    • 1 Introduction 1
    • 2 Background and Related Work 4
    • 2.1 Process Reward Models 4
    • 2.2 Uncertainty Quantification in LLMs 5
    • 2.3 Inference-Time Scaling 6
    • 3 UAPRM 7
    • 3.1 Data Construction 7
    • 3.1.1 Confidence Augmentation via Self-Consistency 8
    • 3.1.2 Addressing Data Imbalance 9
    • 3.2 Training Strategy 9
    • 3.2.1 Token-Wise Weighted Loss 10
    • 3.2.2 Supervised Fine-Tuning with LoRA 10
    • 4 Experiments 14
    • 4.1 Experimental Setup 14
    • 4.2 Verification Performance on ProcessBench 15
    • 4.2.1 Performance Comparison 15
    • 4.2.2 CoT Length Statistics 16
    • 4.3 Best-of-N Scaling on MATH-500 18
    • 4.3.1 BoN Performance 18
    • 4.3.2 Dynamic CoT Adjustment 19
    • 4.3.3 Reliability and Calibration 20
    • 5 Analysis 23
    • 5.1 Mitigating Overconfidence through Weighted Training 23
    • 5.2 Precision-Recall Trade-off in Verification 24
    • 5.3 Limited Dynamic CoT Adjustment 24
    • 5.4 Effectiveness in Best-of-N Selection 25
    • 6 Conclusion 26
    • Bibliography 28
    • A Validation of Label-based Confidence Score 33
    • A.1 Experimental Setup 33
    • A.2 Results 33
    • Abstract (In Korean) 35
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼