오프라인 선호도 기반 강화학습(Offline Preference-based Reinforcement Learning)은 환경과의 상호작용 없이 인간 피드백으로부터 복잡한 행동을 학습하지만, 정책 최적 화 과정에서 분포 외(out-of-distribu...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17450625
서울 : 서울대학교 대학원, 2026
학위논문(석사) -- 서울대학교 대학원 , 협동과정인공지능전공 , 2026. 2
2026
영어
006.3
서울
vi,56 ; 26 cm
지도교수: 조성준
I804:11032-000000194315
0
상세조회0
다운로드오프라인 선호도 기반 강화학습(Offline Preference-based Reinforcement Learning)은 환경과의 상호작용 없이 인간 피드백으로부터 복잡한 행동을 학습하지만, 정책 최적 화 과정에서 분포 외(out-of-distribu...
오프라인 선호도 기반 강화학습(Offline Preference-based Reinforcement Learning)은 환경과의 상호작용 없이 인간 피드백으로부터 복잡한 행동을 학습하지만, 정책 최적
화 과정에서 분포 외(out-of-distribution) 영역을 마주할 때 보상 모델의 외삽 오류 (extrapolation error)를 겪는다. 이러한 오류는 선호도 레이블이 부여된 학습 궤적과
레이블이 없는 추론 데이터 간의 분포 이동(distributional shift)으로부터 발생하며, 보상 오추정과 차선의 정책으로 이어진다. 본 연구는 SPOT(Subgoal-based Preference
Optimization Through Attention Weight)을 제안하며, 이는 선호도 데이터로부터 어텐션 기반으로 도출된 하위 목표를 활용하여 외삽 오류를 완화한다. SPOT은 선호되는
궤적에서 관찰된 하위 목표를 향해 정책을 정규화한다. 이러한 접근법은 학습을 훈련 분포 내로 제약하여 보상 모델의 외삽 오류를 감소시킨다. 포괄적인 실험을 통해, 본
연구의 하위 목표 유도 접근법이 외삽 오류를 줄이면서 기존 방법들 대비 우수한 성능을 달성함을 입증한다. 본 접근법은 세밀한 신용 할당(credit assignment) 정보를 보존하
면서 쿼리 효율성을 향상시키며, 이는 신뢰할 수 있고 실용적인 오프라인 선호도 기반 학습을 위한 유망한 방향을 제시한다.
다국어 초록 (Multilingual Abstract)
Offline preference-based reinforcement learning (PbRL) learns complex behaviors from human feedback without environment interaction, but suffers from reward model extrapolation errors when encountering out-of-distribution region during policy optimiza...
Offline preference-based reinforcement learning (PbRL) learns complex behaviors from human feedback without environment interaction, but suffers from reward
model extrapolation errors when encountering out-of-distribution region during policy optimization. These errors arise from distributional shifts between preference labeled training trajectories and unlabeled inference data, leading to reward misestimation and suboptimal policies. We introduce SPOT (Subgoal-based Preference
Optimization Through Attention Weight), which mitigates extrapolation errors by leveraging attention-derived subgoals from preference data. SPOT regularizes the
policy toward subgoals observed in preferred trajectories. This approach constrains learning within the training distribution, reducing reward model extrapolation errors.
Through comprehensive experiments, we demonstrate that our subgoal-guided approach achieves superior performance compared to existing methods while reducing
extrapolation errors. Our approach preserves fine-grained credit assignment information while enhancing query efficiency, suggesting promising directions for reliable
and practical offline preference-based learning.
목차 (Table of Contents)