RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    A Distributional Perspective on Human-Aligned Decision Making under Uncertainty = 불확실성 속 인간과 정렬된 의사결정에 관한 분포적 접근

    한글로보기

    https://www.riss.kr/link?id=T17450119

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Sequential decision-making under uncertainty is a fundamental problem in artificial intelligence.
    Real-world environments rarely provide well-defined rewards or complete information, and feedback is often qualitative, subjective, or inconsistent.
    As AI systems are increasingly deployed in high-stakes domains such as finance, autonomous driving, and human–robot interaction, it becomes crucial to develop principled algorithms that can act reliably under uncertainty and align with human intentions.
    However, existing reinforcement learning (RL) paradigms, which lack explicit modeling of uncertainty, rely primarily on expectation-based objectives and handcrafted rewards, leaving a substantial gap between theoretical optimality and human-aligned behavior.


    This dissertation addresses these challenges through two complementary perspectives—distributional reinforcement learning (DistRL) and reinforcement learning from human feedback (RLHF)—and unifies them under a common theoretical lens of regret minimization.
    The central goal is to establish a reliable foundation for learning human-aligned decision-making by interpreting the probabilistic nature inherent in human feedback.

    The first part revisits the exploration problem in DistRL.
    Existing approaches based on “optimism under uncertainty” rely on estimates of return variance but conflate epistemic and aleatoric uncertainties, which induces persistent risk-seeking bias and distorted data collection. To address this, we propose the Perturbed Quantile Regression (PQR) algorithm, which introduces randomized perturbations of distorted risk measures to guide action selection.
    We theoretically establish that PQR avoids biased exploration and converges to the true optimum, and empirically show that it outperforms variance-based exploration methods across diverse benchmarks, including 55 Atari games.

    The second part tackles the fundamental challenge of infinite dimensionality in DistRL. Prior work introduced the notion of Bellman closedness, but this fails to guarantee unbiased updates from finite samples in online learning. We propose the concept of Bellman Unbiasedness, which characterizes functionals that are not only preserved under Bellman updates but also estimable without bias from finite samples.
    Our analysis shows that only moment functionals satisfy both conditions. Building on this result, we design the first provably efficient DistRL algorithm under general value function approximation—Statistical Functional Least-Squares Value Iteration (SF-LSVI)—which achieves a tight regret bound of $\tilde{O}(d_E H^{3/2}\sqrt{K})$, improving upon prior results.

    The third part turns to RLHF, where agents learn from preference feedback instead of handcrafted rewards.
    Recent frameworks such as Direct Preference Optimization (DPO) optimize policies directly without an explicit reward model but implicitly assume that all preference data are generated by the optimal policy, leading to a likelihood mismatch.
    To overcome this, we reinterpret preferences through the lens of regret and propose Policy-labeled Preference Learning (PPL), which explicitly integrates policy labels into the learning process.
    Our method introduces contrastive KL regularization that aligns policies with preferred data while contrasting against less-preferred data. We theoretically show that PPL characterizes an equivalence class of reward models consistent with a given optimal policy and establishes statistical robustness via uniquely defined regret. Empirically, PPL substantially improves RLHF performance in offline robotic manipulation tasks and demonstrates robustness in online learning.

    Collectively, these contributions establish regret minimization as a unifying theoretical principle that bridges distributional modeling and human feedback, linking the mathematical efficiency of RL with the behavioral realism of human decision-making.
    This work contributes to the foundation of trustworthy and human-aligned artificial intelligence, providing theoretical and algorithmic insights for robust decision-making under uncertainty.
    번역하기

    Sequential decision-making under uncertainty is a fundamental problem in artificial intelligence. Real-world environments rarely provide well-defined rewards or complete information, and feedback is often qualitative, subjective, or inconsistent. As A...

    Sequential decision-making under uncertainty is a fundamental problem in artificial intelligence.
    Real-world environments rarely provide well-defined rewards or complete information, and feedback is often qualitative, subjective, or inconsistent.
    As AI systems are increasingly deployed in high-stakes domains such as finance, autonomous driving, and human–robot interaction, it becomes crucial to develop principled algorithms that can act reliably under uncertainty and align with human intentions.
    However, existing reinforcement learning (RL) paradigms, which lack explicit modeling of uncertainty, rely primarily on expectation-based objectives and handcrafted rewards, leaving a substantial gap between theoretical optimality and human-aligned behavior.


    This dissertation addresses these challenges through two complementary perspectives—distributional reinforcement learning (DistRL) and reinforcement learning from human feedback (RLHF)—and unifies them under a common theoretical lens of regret minimization.
    The central goal is to establish a reliable foundation for learning human-aligned decision-making by interpreting the probabilistic nature inherent in human feedback.

    The first part revisits the exploration problem in DistRL.
    Existing approaches based on “optimism under uncertainty” rely on estimates of return variance but conflate epistemic and aleatoric uncertainties, which induces persistent risk-seeking bias and distorted data collection. To address this, we propose the Perturbed Quantile Regression (PQR) algorithm, which introduces randomized perturbations of distorted risk measures to guide action selection.
    We theoretically establish that PQR avoids biased exploration and converges to the true optimum, and empirically show that it outperforms variance-based exploration methods across diverse benchmarks, including 55 Atari games.

    The second part tackles the fundamental challenge of infinite dimensionality in DistRL. Prior work introduced the notion of Bellman closedness, but this fails to guarantee unbiased updates from finite samples in online learning. We propose the concept of Bellman Unbiasedness, which characterizes functionals that are not only preserved under Bellman updates but also estimable without bias from finite samples.
    Our analysis shows that only moment functionals satisfy both conditions. Building on this result, we design the first provably efficient DistRL algorithm under general value function approximation—Statistical Functional Least-Squares Value Iteration (SF-LSVI)—which achieves a tight regret bound of $\tilde{O}(d_E H^{3/2}\sqrt{K})$, improving upon prior results.

    The third part turns to RLHF, where agents learn from preference feedback instead of handcrafted rewards.
    Recent frameworks such as Direct Preference Optimization (DPO) optimize policies directly without an explicit reward model but implicitly assume that all preference data are generated by the optimal policy, leading to a likelihood mismatch.
    To overcome this, we reinterpret preferences through the lens of regret and propose Policy-labeled Preference Learning (PPL), which explicitly integrates policy labels into the learning process.
    Our method introduces contrastive KL regularization that aligns policies with preferred data while contrasting against less-preferred data. We theoretically show that PPL characterizes an equivalence class of reward models consistent with a given optimal policy and establishes statistical robustness via uniquely defined regret. Empirically, PPL substantially improves RLHF performance in offline robotic manipulation tasks and demonstrates robustness in online learning.

    Collectively, these contributions establish regret minimization as a unifying theoretical principle that bridges distributional modeling and human feedback, linking the mathematical efficiency of RL with the behavioral realism of human decision-making.
    This work contributes to the foundation of trustworthy and human-aligned artificial intelligence, providing theoretical and algorithmic insights for robust decision-making under uncertainty.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    불확실성 속 인간 피드백 기반 순차적 의사결정을 다루는 본 논문은 분포 강화학습과 인간 피드백 기반 강화학습이라는 두 가지 핵심 연구 분야에 초점을 맞춘다.
    분포 강화학습은 위험에 민감한 제어에서, 인간 피드백 기반 강화학습은 인간 선호도 정렬에서 출발했지만,
    두 분야 모두 불확실성과 불완전한 정보가 필연적인 환경에서 원칙적인 의사결정을 가능하게 하는 알고리즘을 설계해야 한다는 공통된 과제를 공유한다. 본 논문은 이러한 도전 과제들을 후회 최소화라는 통일된 관점에서 해석하고, 에이전트의 행동을 인간의 의사결정 구조에 정렬하기 위한 이론적 프레임워크와 실용적인 알고리즘 원리를 제시한다.

    논문의 첫 번째 부분에서는 분포 강화학습 분야에서의 탐색 문제를 재조명한다. 기존의 `불확실성에 대한 낙관주의'는 반환 분포의 분산 추정치를 활용하지만, 이는 인식론적 불확실성과 내재적 불확실성을 혼동하여 지속적인 위험 추구 편향과 편향된 데이터 수집을 야기하는 문제점이 발생함을 확인하였다. 이를 해결하기 위해 우리는 외란된 퀀타일 정규화 알고리즘(Perturbed Quantile Regression)을 제안한다.
    이는 왜곡된 위험 척도에 무작위 외란이 적용된 척도를 도입하여 행동을 선택하는 방식으로, 이론적으로 편향된 탐색을 피하면서 본래의 최적점에 도달하는 것을 증명하며, 55개의 아타리 게임을 포함한 다양한 벤치마크에서 기존의 분산 기반 탐색 방법보다 우수한 성능을 달성함을 보였다.

    두 번째 부분은 분포 강화학습에서 분포의 무한 차원성이라는 근본적인 난제를 다룬다.
    기존 연구들은 `벨만 닫힘(Bellman closedness)' 개념을 도입했으나, 이는 온라인 학습에서 유한 개의 표본만으로 통계적 함수들이 편향 없이 업데이트될 수 있음을 보장하지 못하는 한계가 존재한다.
    이에 우리는 벨만 업데이트에서 보존될 뿐만 아니라 유한 개의 샘플로부터 편향 없이 추정 가능한 기능을 특징짓는 '벨만 비편향성(Bellman Unbiasedness)' 개념을 제안한다.
    우리의 분석은 오직 모멘트 함수족만이 이 두 가지 특성을 만족함을 밝히고, 이를 바탕으로 일반적인 가치 함수 근사에서도 이론적으로 효율성을 갖춘 최초의 분포형 강화학습 알고리즘인 `통계적 함수 기반 최소제곱 가치 반복 알고리즘(Statistical Functional Least-Squares Value Iteration)'을 설계하였다.
    이는 이전 연구들보다 향상된 $\tilde{O}(d_E H^{3/2}\sqrt{K})$라는 타이트한 후회 상한선을 달성한다.

    논문의 세 번째 부분은 인간의 수작업 보상 대신 선호도 피드백으로부터 학습하는 인간 피드백 기반 강화학습을 다룬다.
    직접 선호도 최적화(Direct Preference Optimization)와 같은 최근 프레임워크는 보상 모델 없이 정책을 직접 최적화하지만, 모든 데이터가 최적의 정책에 의해 생성되었다고 가정하는 '우도 불일치(likelihood mismatch)' 문제를 내재적으로 전제하고 있음을 밝힌다.
    이를 해결하기 위해 우리는 후회 개념을 활용하여 인간 선호도를 재해석하고 행동 정책 레이블을 학습 과정에 명시적으로 통합하는 `정책 레이블 기반 선호학습(Policy-labeled Preference Learning)'을 제안한다. 제안하는 알고리즘은 선호되는 데이터에 정책을 맞추고 덜 선호되는 데이터와 대조하는 '대조적 KL 정규화'를 도입한다.
    이론적으로 주어진 최적 정책에 대해 보상체계의 등가 클래스를 제공하며, 후회가 유일하게 정의됨에 따른 통계적 강건성을 입증하였다. 실험적으로 로봇 조작 작업에서 오프라인 학습 환경에서의 인간 피드백 기반 강화학습의 성능을 크게 향상시키고 온라인 학습 환경에서 강건함을 입증하였다.

    요약하자면, 본 논문은 (1) 분포 강화학습의 편향된 탐색 문제를 해결하는 알고리즘, (2) 일반적인 가치 함수 근사에서 온라인 분포 업데이트를 위한 최초의 증명 가능한 효율적 프레임워크를 제공하는 벨만 비편향성 개념과 이에 기반한 알고리즘, (3) 후회 기반 선호도 모델링을 통해 인간 피드백 기반 강화학습의 우도 불일치 문제를 해결하는 알고리즘을 제시한다. 이러한 연구들은 후회 최소화를 이론과 실무를 아우르는 통일된 원리로 정립하며, 불확실성 속에서도 신뢰성 높고 인간과 정렬된 인공지능 의사결정을 가능하게 하는 기반을 제공한다.
    번역하기

    불확실성 속 인간 피드백 기반 순차적 의사결정을 다루는 본 논문은 분포 강화학습과 인간 피드백 기반 강화학습이라는 두 가지 핵심 연구 분야에 초점을 맞춘다. 분포 강화학습은 위험에 ...

    불확실성 속 인간 피드백 기반 순차적 의사결정을 다루는 본 논문은 분포 강화학습과 인간 피드백 기반 강화학습이라는 두 가지 핵심 연구 분야에 초점을 맞춘다.
    분포 강화학습은 위험에 민감한 제어에서, 인간 피드백 기반 강화학습은 인간 선호도 정렬에서 출발했지만,
    두 분야 모두 불확실성과 불완전한 정보가 필연적인 환경에서 원칙적인 의사결정을 가능하게 하는 알고리즘을 설계해야 한다는 공통된 과제를 공유한다. 본 논문은 이러한 도전 과제들을 후회 최소화라는 통일된 관점에서 해석하고, 에이전트의 행동을 인간의 의사결정 구조에 정렬하기 위한 이론적 프레임워크와 실용적인 알고리즘 원리를 제시한다.

    논문의 첫 번째 부분에서는 분포 강화학습 분야에서의 탐색 문제를 재조명한다. 기존의 `불확실성에 대한 낙관주의'는 반환 분포의 분산 추정치를 활용하지만, 이는 인식론적 불확실성과 내재적 불확실성을 혼동하여 지속적인 위험 추구 편향과 편향된 데이터 수집을 야기하는 문제점이 발생함을 확인하였다. 이를 해결하기 위해 우리는 외란된 퀀타일 정규화 알고리즘(Perturbed Quantile Regression)을 제안한다.
    이는 왜곡된 위험 척도에 무작위 외란이 적용된 척도를 도입하여 행동을 선택하는 방식으로, 이론적으로 편향된 탐색을 피하면서 본래의 최적점에 도달하는 것을 증명하며, 55개의 아타리 게임을 포함한 다양한 벤치마크에서 기존의 분산 기반 탐색 방법보다 우수한 성능을 달성함을 보였다.

    두 번째 부분은 분포 강화학습에서 분포의 무한 차원성이라는 근본적인 난제를 다룬다.
    기존 연구들은 `벨만 닫힘(Bellman closedness)' 개념을 도입했으나, 이는 온라인 학습에서 유한 개의 표본만으로 통계적 함수들이 편향 없이 업데이트될 수 있음을 보장하지 못하는 한계가 존재한다.
    이에 우리는 벨만 업데이트에서 보존될 뿐만 아니라 유한 개의 샘플로부터 편향 없이 추정 가능한 기능을 특징짓는 '벨만 비편향성(Bellman Unbiasedness)' 개념을 제안한다.
    우리의 분석은 오직 모멘트 함수족만이 이 두 가지 특성을 만족함을 밝히고, 이를 바탕으로 일반적인 가치 함수 근사에서도 이론적으로 효율성을 갖춘 최초의 분포형 강화학습 알고리즘인 `통계적 함수 기반 최소제곱 가치 반복 알고리즘(Statistical Functional Least-Squares Value Iteration)'을 설계하였다.
    이는 이전 연구들보다 향상된 $\tilde{O}(d_E H^{3/2}\sqrt{K})$라는 타이트한 후회 상한선을 달성한다.

    논문의 세 번째 부분은 인간의 수작업 보상 대신 선호도 피드백으로부터 학습하는 인간 피드백 기반 강화학습을 다룬다.
    직접 선호도 최적화(Direct Preference Optimization)와 같은 최근 프레임워크는 보상 모델 없이 정책을 직접 최적화하지만, 모든 데이터가 최적의 정책에 의해 생성되었다고 가정하는 '우도 불일치(likelihood mismatch)' 문제를 내재적으로 전제하고 있음을 밝힌다.
    이를 해결하기 위해 우리는 후회 개념을 활용하여 인간 선호도를 재해석하고 행동 정책 레이블을 학습 과정에 명시적으로 통합하는 `정책 레이블 기반 선호학습(Policy-labeled Preference Learning)'을 제안한다. 제안하는 알고리즘은 선호되는 데이터에 정책을 맞추고 덜 선호되는 데이터와 대조하는 '대조적 KL 정규화'를 도입한다.
    이론적으로 주어진 최적 정책에 대해 보상체계의 등가 클래스를 제공하며, 후회가 유일하게 정의됨에 따른 통계적 강건성을 입증하였다. 실험적으로 로봇 조작 작업에서 오프라인 학습 환경에서의 인간 피드백 기반 강화학습의 성능을 크게 향상시키고 온라인 학습 환경에서 강건함을 입증하였다.

    요약하자면, 본 논문은 (1) 분포 강화학습의 편향된 탐색 문제를 해결하는 알고리즘, (2) 일반적인 가치 함수 근사에서 온라인 분포 업데이트를 위한 최초의 증명 가능한 효율적 프레임워크를 제공하는 벨만 비편향성 개념과 이에 기반한 알고리즘, (3) 후회 기반 선호도 모델링을 통해 인간 피드백 기반 강화학습의 우도 불일치 문제를 해결하는 알고리즘을 제시한다. 이러한 연구들은 후회 최소화를 이론과 실무를 아우르는 통일된 원리로 정립하며, 불확실성 속에서도 신뢰성 높고 인간과 정렬된 인공지능 의사결정을 가능하게 하는 기반을 제공한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Chapter 1. Introduction 1
    • 1.1 The Shift from Perception to Action 1
    • 1.2 The Cognitive Gap: Uncertainty and Regret in Human Decision-Making 3
    • Abstract i
    • Chapter 1. Introduction 1
    • 1.1 The Shift from Perception to Action 1
    • 1.2 The Cognitive Gap: Uncertainty and Regret in Human Decision-Making 3
    • 1.3 Research Scope and Unified Hypothesis on Uncertainty 6
    • 1.4 Core Research Areas and Contributions 8
    • 1.4.1 Distributional Reinforcement Learning: Uncertainty, Efficiency, and Bias 8
    • 1.4.2 Reinforcement Learning from Human Feedback: Robust Alignment via Regret 10
    • 1.5 Organization of the Dissertation 11
    • 1.6 Publications 12
    • Chapter 2. Background 13
    • 2.1 Reinforcement Learning and Markov Decision Processes 14
    • 2.1.1 Markov Decision Processes 14
    • 2.1.2 Limitations of Classical RL 15
    • 2.2 Distributional Reinforcement Learning 17
    • 2.2.1 Distributional Bellman Equation and Convergence Properties 17
    • 2.2.2 Distributional Bellman Optimality and Instabilities 18
    • 2.2.3 Approximation Schemes in Distributional RL 19
    • 2.3 Reinforcement Learning from Human Feedback 23
    • 2.3.1 Motivation and Origins 23
    • 2.3.2 Canonical RLHF Pipeline 23
    • 2.3.3 Direct Preference Optimization (DPO) and Extensions 25
    • 2.3.4 Challenges and Open Directions 26
    • 2.4 Regret Minimization Framework 28
    • 2.4.1 From Expected Utility to Prospect Theory 28
    • 2.4.2 Regret Theory: Anticipating Counterfactual Emotion 29
    • 2.4.3 Algorithmic Regret in Reinforcement Learning 31
    • 2.4.4 Bridging Behavioral and Algorithmic Perspectives 31
    • 2.4.5 Toward a Unified View 32
    • 2.5 Summary 33
    • Chapter 3. Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk Criterion 34
    • 3.1 Backgrounds & Related Works 37
    • 3.1.1 Distributional RL 37
    • 3.1.2 Exploration on Distributional RL 38
    • 3.1.3 Risk in Distributional RL 39
    • 3.2 Perturbation in Distributional RL 40
    • 3.2.1 Perturbed Distributional Bellman Optimality Operator 40
    • 3.2.2 Convergence of the Perturbed Distributional Bellman Optimality Operator 43
    • 3.2.3 Practical Algorithm with Distributional Perturbation 44
    • 3.3 Experiments on Stochastic Environments with High Intrinsic Uncertainty 46
    • 3.3.1 N-Chain Environment 47
    • 3.3.2 LunarLander-v2 52
    • 3.3.3 55 Atari Games 53
    • 3.4 Related Works & Discussion 59
    • 3.4.1 Comparison with QUOTA 59
    • 3.4.2 Reproducibility Issues on DLTV 61
    • 3.5 Summary 62
    • Chapter 4. Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation 64
    • 4.1 Related Work 66
    • 4.2 Preliminaries 68
    • 4.3 Statistical Functionals in Distributional RL 70
    • 4.3.1 Bellman Closedness 71
    • 4.3.2 Bellman Unbiasedness 73
    • 4.3.3 Statistical Functional Bellman Completeness 76
    • 4.4 SF-LSVI: Statistical Functional Least Squares Value Iteration 78
    • 4.5 Theoretical Analysis 79
    • 4.6 Summary 82
    • Chapter 5. Policy-labeled Preference Learning: Is Preference Enough for RLHF? 84
    • 5.1 Preliminaries 85
    • 5.1.1 Preference-based Reinforcement Learning 86
    • 5.2 Policy-labeled Preference Learning 89
    • 5.2.1 Is Preference Enough for RLHF? 90
    • 5.2.2 Theoretical Analysis 93
    • 5.2.3 Practical Algorithm and Implementation Details 99
    • 5.3 Experiments 102
    • 5.3.1 Experimental Setup 102
    • 5.3.2 Can PPL be effectively trained on both homogeneous and heterogeneous offline datasets? 103
    • 5.3.3 Does incorporating policy labels improve learning performance? 105
    • 5.3.4 Can PPL be effectively applied to an online RLHF algorithm? 106
    • 5.4 Summary 107
    • Chapter 6. Conclusion 108
    • 6.1 Future Work 109
    • Appendix A. Appendix of Chapter 3 111
    • A.1 Main Proof 111
    • A.1.1 Technical Lemma 111
    • A.1.2 Proof of Theorem A.1.3 112
    • A.1.3 Proof of Theorem 3.2.3 113
    • A.1.4 Proof of Theorem 3.2.4 115
    • A.2 Implementation Details 117
    • A.2.1 Hyperparameter Setting 117
    • A.3 Raw Scores across 55 Atari Games 118
    • Appendix B. Appendix of Chapter 4 121
    • B.1 Notation 121
    • B.2 Pseudocode of SF-LSVI and Technical Remarks 124
    • B.3 Related Work and Discussion 125
    • B.3.1 Technical Clarifications on Linearity Assumption in Existing Results 125
    • B.3.2 Existence of Nonlinear Bellman Closed Sketch 126
    • B.3.3 Non-existence of Sketch Bellman Operator for Quantile Functional 127
    • B.4 Proof 131
    • Appendix C. Appendix of Chapter 5 144
    • C.1 Main Proof 144
    • C.2 Further Theoretical Analysis & Discussion 147
    • C.2.1 Mathematical Derivation of PPL Framework 147
    • C.3 Variants of PPL and Baselines 149
    • C.4 Implementation Details 151
    • C.4.1 Hyperparameter Setting 151
    • C.4.2 MetaWorld Benchmark 152
    • C.4.3 Reproducibility Check 153
    • C.4.4 Offline Dataset Generation and Its Distribution 156
    • C.4.5 Online Implementation 158
    • C.5 Experimental Results on Homogeneous / Heterogeneous Datasets 160
    • C.6 Comparison with Deterministic Pseudo-labels 164
    • C.7 Experimental Results on Online Implementation 166
    • 초록 185
    • 감사의 글 187
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼