RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Mitigating Length Bias in RLHF through a Causal Lens = 인과적 접근을 통한 RLHF의 길이 편향 완화

    한글로보기

    https://www.riss.kr/link?id=T17315390

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Reinforcement learning from human feedback (RLHF) has become a standard technique for aligning large language models (LLMs) with human intentions. However, reward models trained under this paradigm often exhibit a consistent bias toward longer responses, mistakenly associating verbosity with higher quality. To address this issue, we introduce a causal intervention framework that disentangles content informativeness from response length. At the core of our method is a counterfactual data augmentation strategy that creates response pairs with controlled variation: (1) pairs differing in length but conveying the same information, and (2) pairs with varied semantic content but matched in length. These examples allow the reward model to focus on content quality while discounting superficial verbosity. Experimental results confirm that our approach effectively reduces length bias in reward scoring and produces more concise and semantically focused policy outputs. Our findings suggest that causal data augmentation can enhance the robustness and content sensitivity of reward modeling in RLHF pipelines.
    번역하기

    Reinforcement learning from human feedback (RLHF) has become a standard technique for aligning large language models (LLMs) with human intentions. However, reward models trained under this paradigm often exhibit a consistent bias toward longer respons...

    Reinforcement learning from human feedback (RLHF) has become a standard technique for aligning large language models (LLMs) with human intentions. However, reward models trained under this paradigm often exhibit a consistent bias toward longer responses, mistakenly associating verbosity with higher quality. To address this issue, we introduce a causal intervention framework that disentangles content informativeness from response length. At the core of our method is a counterfactual data augmentation strategy that creates response pairs with controlled variation: (1) pairs differing in length but conveying the same information, and (2) pairs with varied semantic content but matched in length. These examples allow the reward model to focus on content quality while discounting superficial verbosity. Experimental results confirm that our approach effectively reduces length bias in reward scoring and produces more concise and semantically focused policy outputs. Our findings suggest that causal data augmentation can enhance the robustness and content sensitivity of reward modeling in RLHF pipelines.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    인간 피드백을 활용한 강화학습(RLHF)은 대형 언어 모델(LLM)의 출력을 인간의 의도에 맞게 조정하기 위한 대표적인 기법으로 자리 잡고 있다. 그러나 이 방식으로 학습된 보상 모델은 종종 응답의 품질과 장황함을 혼동하여, 길이가 긴 응답에 일관되게 높은 보상을 부여하는 \textit{길이 편향(length bias)}을 나타낸다. 본 연구에서는 이러한 편향을 분석하고 완화하기 위한 인과적(interventional) 프레임워크를 제안한다. 핵심 아이디어는, 정보량은 같지만 길이가 다른 응답 쌍, 혹은 길이는 유사하지만 내용이 다른 응답 쌍을 생성하는 \textit{반사실적 데이터 증강(counterfactual data augmentation)} 기법을 활용하는 것이다. 이러한 제어된 비교쌍을 통해 보상 모델은 장황함이 아닌 내용 품질에 기반하여 응답을 평가할 수 있도록 유도된다. 실험 결과, 제안하는 방법은 보상 모델의 길이 편향을 효과적으로 줄이며, 정책 모델로부터 더 간결하고 의미 중심적인 출력을 유도함을 확인하였다. 본 연구는 인과 기반 데이터 증강이 RLHF 보상 학습의 견고성과 내용 민감도를 향상시킬 수 있음을 시사한다.
    번역하기

    인간 피드백을 활용한 강화학습(RLHF)은 대형 언어 모델(LLM)의 출력을 인간의 의도에 맞게 조정하기 위한 대표적인 기법으로 자리 잡고 있다. 그러나 이 방식으로 학습된 보상 모델은 종종 응...

    인간 피드백을 활용한 강화학습(RLHF)은 대형 언어 모델(LLM)의 출력을 인간의 의도에 맞게 조정하기 위한 대표적인 기법으로 자리 잡고 있다. 그러나 이 방식으로 학습된 보상 모델은 종종 응답의 품질과 장황함을 혼동하여, 길이가 긴 응답에 일관되게 높은 보상을 부여하는 \textit{길이 편향(length bias)}을 나타낸다. 본 연구에서는 이러한 편향을 분석하고 완화하기 위한 인과적(interventional) 프레임워크를 제안한다. 핵심 아이디어는, 정보량은 같지만 길이가 다른 응답 쌍, 혹은 길이는 유사하지만 내용이 다른 응답 쌍을 생성하는 \textit{반사실적 데이터 증강(counterfactual data augmentation)} 기법을 활용하는 것이다. 이러한 제어된 비교쌍을 통해 보상 모델은 장황함이 아닌 내용 품질에 기반하여 응답을 평가할 수 있도록 유도된다. 실험 결과, 제안하는 방법은 보상 모델의 길이 편향을 효과적으로 줄이며, 정책 모델로부터 더 간결하고 의미 중심적인 출력을 유도함을 확인하였다. 본 연구는 인과 기반 데이터 증강이 RLHF 보상 학습의 견고성과 내용 민감도를 향상시킬 수 있음을 시사한다.

    더보기

    목차 (Table of Contents)

    • Chapter 1 Introduction 1
    • Chapter 2 Preliminaries 4
    • 2.0.1 Reinforcement Learning from Human Feedback 4
    • 2.0.2 Reward Hacking and Length Bias 6
    • Chapter 1 Introduction 1
    • Chapter 2 Preliminaries 4
    • 2.0.1 Reinforcement Learning from Human Feedback 4
    • 2.0.2 Reward Hacking and Length Bias 6
    • 2.0.3 Existing Approaches on Length Bias 7
    • 2.0.4 Understanding of Counterfactuals 10
    • Chapter 3 Causal Interpretation of Length Bias 16
    • 3.0.1 Length Bias as a Causal Problem 16
    • 3.0.2 Motivation for Counterfactuals 18
    • 3.0.3 Feasibility of Counterfactuals 21
    • 3.0.4 Operational Definitions of Length and Content 28
    • Chapter 4 Length Bias Mitigation Pipeline 29
    • 4.0.1 Counterfactual Data Augmentation Implementations 29
    • 4.0.2 Diagnosing Length Bias 31
    • 4.0.3 Mitigating Length Bias 33
    • Chapter 5 Experiments 35
    • 5.0.1 Data Augmentation 35
    • 5.0.2 Length Bias Identification and Mitigation Data Construction 38
    • Chapter 6 Conclusion 50
    • 요약 60
    • Appendix A Motivating Details 61
    • A.1 Choices of Odin and RRM 61
    • A.2 Response Curve 67
    • Appendix B Experimental Details 70
    • B.1 Augmentation Techniques for Fixing and Varying Content 70
    • B.2 Augmentation Examples 73
    • B.3 Fine-Tuning Details 73
    • B.3.1 Cross-Encoder Fine-Tuning 73
    • B.3.2 Reward Model Fine-Tuning Details 75
    • B.3.3 Supervised Fine-Tuning (SFT) Details 77
    • B.3.4 ODIN Fine-Tuning Details 79
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼