RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Mitigating Reward Extrapolation Errors with Subgoal-based Preference Optimization through Attention Weight in Offline Preference-Based Reinforcement Learning = 어텐션 기반 하위 목표 발견을 통한 오프라인 선호도 기반 강화학습에서의 보상 외삽 오류 완화

    한글로보기

    https://www.riss.kr/link?id=T17450625

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    오프라인 선호도 기반 강화학습(Offline Preference-based Reinforcement Learning)은 환경과의 상호작용 없이 인간 피드백으로부터 복잡한 행동을 학습하지만, 정책 최적
    화 과정에서 분포 외(out-of-distribution) 영역을 마주할 때 보상 모델의 외삽 오류 (extrapolation error)를 겪는다. 이러한 오류는 선호도 레이블이 부여된 학습 궤적과
    레이블이 없는 추론 데이터 간의 분포 이동(distributional shift)으로부터 발생하며, 보상 오추정과 차선의 정책으로 이어진다. 본 연구는 SPOT(Subgoal-based Preference
    Optimization Through Attention Weight)을 제안하며, 이는 선호도 데이터로부터 어텐션 기반으로 도출된 하위 목표를 활용하여 외삽 오류를 완화한다. SPOT은 선호되는
    궤적에서 관찰된 하위 목표를 향해 정책을 정규화한다. 이러한 접근법은 학습을 훈련 분포 내로 제약하여 보상 모델의 외삽 오류를 감소시킨다. 포괄적인 실험을 통해, 본
    연구의 하위 목표 유도 접근법이 외삽 오류를 줄이면서 기존 방법들 대비 우수한 성능을 달성함을 입증한다. 본 접근법은 세밀한 신용 할당(credit assignment) 정보를 보존하
    면서 쿼리 효율성을 향상시키며, 이는 신뢰할 수 있고 실용적인 오프라인 선호도 기반 학습을 위한 유망한 방향을 제시한다.
    번역하기

    오프라인 선호도 기반 강화학습(Offline Preference-based Reinforcement Learning)은 환경과의 상호작용 없이 인간 피드백으로부터 복잡한 행동을 학습하지만, 정책 최적 화 과정에서 분포 외(out-of-distribu...

    오프라인 선호도 기반 강화학습(Offline Preference-based Reinforcement Learning)은 환경과의 상호작용 없이 인간 피드백으로부터 복잡한 행동을 학습하지만, 정책 최적
    화 과정에서 분포 외(out-of-distribution) 영역을 마주할 때 보상 모델의 외삽 오류 (extrapolation error)를 겪는다. 이러한 오류는 선호도 레이블이 부여된 학습 궤적과
    레이블이 없는 추론 데이터 간의 분포 이동(distributional shift)으로부터 발생하며, 보상 오추정과 차선의 정책으로 이어진다. 본 연구는 SPOT(Subgoal-based Preference
    Optimization Through Attention Weight)을 제안하며, 이는 선호도 데이터로부터 어텐션 기반으로 도출된 하위 목표를 활용하여 외삽 오류를 완화한다. SPOT은 선호되는
    궤적에서 관찰된 하위 목표를 향해 정책을 정규화한다. 이러한 접근법은 학습을 훈련 분포 내로 제약하여 보상 모델의 외삽 오류를 감소시킨다. 포괄적인 실험을 통해, 본
    연구의 하위 목표 유도 접근법이 외삽 오류를 줄이면서 기존 방법들 대비 우수한 성능을 달성함을 입증한다. 본 접근법은 세밀한 신용 할당(credit assignment) 정보를 보존하
    면서 쿼리 효율성을 향상시키며, 이는 신뢰할 수 있고 실용적인 오프라인 선호도 기반 학습을 위한 유망한 방향을 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Offline preference-based reinforcement learning (PbRL) learns complex behaviors from human feedback without environment interaction, but suffers from reward
    model extrapolation errors when encountering out-of-distribution region during policy optimization. These errors arise from distributional shifts between preference labeled training trajectories and unlabeled inference data, leading to reward misestimation and suboptimal policies. We introduce SPOT (Subgoal-based Preference
    Optimization Through Attention Weight), which mitigates extrapolation errors by leveraging attention-derived subgoals from preference data. SPOT regularizes the
    policy toward subgoals observed in preferred trajectories. This approach constrains learning within the training distribution, reducing reward model extrapolation errors.
    Through comprehensive experiments, we demonstrate that our subgoal-guided approach achieves superior performance compared to existing methods while reducing
    extrapolation errors. Our approach preserves fine-grained credit assignment information while enhancing query efficiency, suggesting promising directions for reliable
    and practical offline preference-based learning.
    번역하기

    Offline preference-based reinforcement learning (PbRL) learns complex behaviors from human feedback without environment interaction, but suffers from reward model extrapolation errors when encountering out-of-distribution region during policy optimiza...

    Offline preference-based reinforcement learning (PbRL) learns complex behaviors from human feedback without environment interaction, but suffers from reward
    model extrapolation errors when encountering out-of-distribution region during policy optimization. These errors arise from distributional shifts between preference labeled training trajectories and unlabeled inference data, leading to reward misestimation and suboptimal policies. We introduce SPOT (Subgoal-based Preference
    Optimization Through Attention Weight), which mitigates extrapolation errors by leveraging attention-derived subgoals from preference data. SPOT regularizes the
    policy toward subgoals observed in preferred trajectories. This approach constrains learning within the training distribution, reducing reward model extrapolation errors.
    Through comprehensive experiments, we demonstrate that our subgoal-guided approach achieves superior performance compared to existing methods while reducing
    extrapolation errors. Our approach preserves fine-grained credit assignment information while enhancing query efficiency, suggesting promising directions for reliable
    and practical offline preference-based learning.

    더보기

    목차 (Table of Contents)

    • Contents
    • Abstract i
    • Contents iv
    • List of Tables v
    • List of Figures vi
    • Contents
    • Abstract i
    • Contents iv
    • List of Tables v
    • List of Figures vi
    • Chapter 1 Introduction 1
    • 1.1 Problem Description 1
    • 1.2 Research Motivation and Contribution 3
    • Chapter 2 Related Work 5
    • 2.1 Offline Preference-based Reinforcement Learning 5
    • 2.2 Extrapolation Error 6
    • 2.3 Conditional Variational Autoencoders 7
    • 2.4 Offline PbRL Markov Decision Process (MDP) 8
    • 2.5 Preference Transformer 10
    • 2.6 Baseline Comparison 11
    • Chapter 3 Proposed Method 12
    • 3.1 Subgoal Learning via CVAE 12
    • 3.1.1 Attention-Based Subgoal Identification 12
    • 3.1.2 Dual-criteria filtering 13
    • 3.1.3 Conditional Variational Autoencoder Training 14
    • 3.2 Reward Shaping for Offline RL 15
    • 3.2.1 Sub-goal-Guided Reward Augmentation 15
    • 3.2.2 Integrated Reward Signal 16
    • 3.3 Pseudo Code 17
    • Chapter 4 Experiments 19
    • 4.1 Experiments Settings 19
    • 4.1.1 Datasets and Tasks Detail 19
    • 4.1.2 Experimental Details 21
    • 4.2 Benchmark Result 24
    • 4.3 Extrapolation Error Analysis in SPOT 27
    • 4.4 Analysis of Top-K% Subgoal Performance 29
    • 4.5 Analysis of Reward Shaping Methods and Weight Selection 30
    • 4.6 Subgoal Extraction Case Study 30
    • 4.7 Query Efficiency 32
    • Chapter 5 Ablation studies 33
    • 5.1 Auxiliary Loss Functions 33
    • 5.2 Computation Time 33
    • 5.3 Subgoal–Observation Alignment 35
    • 5.4 Subgoal quality metrics 36
    • 5.5 Subgoal Extraction in Goal-Oriented Manipulation Tasks 37
    • Chapter 6 Conclusion 40
    • 6.1 Summary 40
    • 6.2 Limitation & Future work 40
    • Bibliography 42
    • 국문초록 55
    • 감사의 글 56
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼