RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Exact Prompt-Wise Trust-Region Optimization for Fine-Tuning LLMs with Reinforcement Learning = 강화학습 기반 대규모 언어 모델 미세조정을 위한 정확한 프롬프트 단위 신뢰영역 최적화

    한글로보기

    https://www.riss.kr/link?id=T17450286

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Reinforcement learning (RL) has emerged as a core approach for enhancing the reasoning abilities of large language models (LLMs). However, existing RL methods often suffer from premature entropy collapse, selective over-updating of high-reward trajectories, and instability arising from poorly controlled optimization dynamics. Crucially, current approaches lack a principled mechanism for constraining policy updates at the prompt level, which leads to inconsistent learning behavior across trajectories. We introduce Prompt-wise Trust-Region Optimization (PTRO), a new algorithm that formulates RL finetuning as an exact prompt-wise trust-region optimization problem. This formulation yields a principled dual update that explicitly regulates per-prompt update magnitude and removes the need for heuristic clipping or global KL penalties. PTRO achieves higher Pass@k, slower entropy decay, and significantly more stable training, remaining robust even under larger learning rates and deeper inner-loop updates. Our results demonstrate that prompt-wise trust-region optimization provides a simple, principled, and highly effective foundation for stable RL fine-tuning of LLMs.
    번역하기

    Reinforcement learning (RL) has emerged as a core approach for enhancing the reasoning abilities of large language models (LLMs). However, existing RL methods often suffer from premature entropy collapse, selective over-updating of high-reward traject...

    Reinforcement learning (RL) has emerged as a core approach for enhancing the reasoning abilities of large language models (LLMs). However, existing RL methods often suffer from premature entropy collapse, selective over-updating of high-reward trajectories, and instability arising from poorly controlled optimization dynamics. Crucially, current approaches lack a principled mechanism for constraining policy updates at the prompt level, which leads to inconsistent learning behavior across trajectories. We introduce Prompt-wise Trust-Region Optimization (PTRO), a new algorithm that formulates RL finetuning as an exact prompt-wise trust-region optimization problem. This formulation yields a principled dual update that explicitly regulates per-prompt update magnitude and removes the need for heuristic clipping or global KL penalties. PTRO achieves higher Pass@k, slower entropy decay, and significantly more stable training, remaining robust even under larger learning rates and deeper inner-loop updates. Our results demonstrate that prompt-wise trust-region optimization provides a simple, principled, and highly effective foundation for stable RL fine-tuning of LLMs.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    강화학습은 대규모 언어 모델의 추론 능력을 향상시키는 데 있어 핵심적인 방법으로 널리 활용되고 있다. 그러나 기존의 강화학습 기반 미세조정 기법들은 학습 초기에 엔트로피가 빠르게 붕괴되거나, 일부 높은 보상을 받은 궤적에 대해 과도한 업데이트가 집중되는 현상, 최적화 과정이 충분히 제어되지 않아 학습이 불안정한 문제를 공통적으로 겪고 있다. 특히, 현재 방법들에는 프롬프트 단위에서 정책 업데이트의 크기를 원리적으로 제어할 수 있는 메커니즘이 없다. 본 연구에서는 이러한 문제를 해결하기 위해, 강화학습 기반 LLM 미세조정을 프롬프트 단위 신뢰영역 최적화 문제로 정확하게 정식화한 새로운 알고리즘인 Prompt-wise Trust-Region Optimization (PTRO)을 제안한다. 제안한 방법은 각 프롬프트에 대해 허용 가능한 정책 변화의 범위를 명시적으로 제한하는 이중 최적화 규칙을 도출하며, 기존 방법에서 널리 사용되던 휴리스틱한 클리핑 없이도 안정적인 학습을 가능하게 한다. 다양한 수학 추론 벤치마크 실험을 통해, PTRO는 Pass@k 성능을 지속적으로 향상시키는 동시에 엔트로피 감소를 완만하게 유지하며, 더 큰 학습률이나 깊은 내부 업데이트 설정에서도 안정적인 학습 특성을 보임을 확인하였다. 이러한 결과는 프롬프트 단위 신뢰영역 최적화가 대규모 언어 모델의 강화학습 미세조정을 위한 단순하면서도 원리적인 대안이 될 수 있음을 시사한다.
    번역하기

    강화학습은 대규모 언어 모델의 추론 능력을 향상시키는 데 있어 핵심적인 방법으로 널리 활용되고 있다. 그러나 기존의 강화학습 기반 미세조정 기법들은 학습 초기에 엔트로피가 빠르게 ...

    강화학습은 대규모 언어 모델의 추론 능력을 향상시키는 데 있어 핵심적인 방법으로 널리 활용되고 있다. 그러나 기존의 강화학습 기반 미세조정 기법들은 학습 초기에 엔트로피가 빠르게 붕괴되거나, 일부 높은 보상을 받은 궤적에 대해 과도한 업데이트가 집중되는 현상, 최적화 과정이 충분히 제어되지 않아 학습이 불안정한 문제를 공통적으로 겪고 있다. 특히, 현재 방법들에는 프롬프트 단위에서 정책 업데이트의 크기를 원리적으로 제어할 수 있는 메커니즘이 없다. 본 연구에서는 이러한 문제를 해결하기 위해, 강화학습 기반 LLM 미세조정을 프롬프트 단위 신뢰영역 최적화 문제로 정확하게 정식화한 새로운 알고리즘인 Prompt-wise Trust-Region Optimization (PTRO)을 제안한다. 제안한 방법은 각 프롬프트에 대해 허용 가능한 정책 변화의 범위를 명시적으로 제한하는 이중 최적화 규칙을 도출하며, 기존 방법에서 널리 사용되던 휴리스틱한 클리핑 없이도 안정적인 학습을 가능하게 한다. 다양한 수학 추론 벤치마크 실험을 통해, PTRO는 Pass@k 성능을 지속적으로 향상시키는 동시에 엔트로피 감소를 완만하게 유지하며, 더 큰 학습률이나 깊은 내부 업데이트 설정에서도 안정적인 학습 특성을 보임을 확인하였다. 이러한 결과는 프롬프트 단위 신뢰영역 최적화가 대규모 언어 모델의 강화학습 미세조정을 위한 단순하면서도 원리적인 대안이 될 수 있음을 시사한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • 1 Introduction 1
    • 2 Preliminaries 5
    • 2.1 KL-Constrained Trust-Region Optimization 6
    • Abstract i
    • Contents ii
    • 1 Introduction 1
    • 2 Preliminaries 5
    • 2.1 KL-Constrained Trust-Region Optimization 6
    • 2.2 Limitations of Existing Trust-Region Methods 8
    • 3 Methods 11
    • 3.1 Prompt-wise Trust-Region Policy Optimization 12
    • 3.2 Objective Function and Algorithm 14
    • 3.3 Interpretation and Key Observations 18
    • 4 Experiments 20
    • 4.1 Experimental Setup 20
    • 4.2 Results and Analysis 22
    • 4.3 Stability Analysis 25
    • 5 Related Work 27
    • 6 Conclusion 29
    • 7 Appendix: Proof 31
    • 7.1 Derivation of the Exact Query-wise Trust-Region Policy Update 31
    • 7.2 Deriving Optimal Dual Variable 34
    • 7.3 Loss Derivation 35
    • 7.4 Proof of Theorem 3.5 36
    • Abstract (In Korean) 46
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼