RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Resilience of Reinforcement Learning in Out-of-Distribution = 학습 분포 외 상황에서 강화학습의 복원력에 관한 연구

    한글로보기

    https://www.riss.kr/link?id=T17315142

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 논문은 강화학습 에이전트의 학습 분포 외 상황에서의 복원력 향상 문제를 다룬다. 심층 강화학습은 다양한 분야에서 뛰어난 성과를 거두었지만, 학습 분포를 벗어난 상황에서는 신뢰할 수 없는 행동을 생성해 임무 수행의 실패로 이어질 수 있다. 기존 연구들은 주로 학습 분포를 확장하거나 에이전트가 학습 분포 외 상황에 처하지 않도록 예방하는 데 초점을 맞추어 왔으나, 실제로 학습 분포 외 상황에 직면했을 때 에이전트가 어떻게 대응해야 하는지에 대해서는 충분히 다루지 않았다.

    이를 해결하기 위해 본 논문에서는 학습 분포 외 상황에서 강화학습 에이전트의 복원력을 정의하고, 이를 향상시키기 위한 새로운 방법론들을 제안한다. 3장에서는 불확실성 기반 조건부 정책을 도입하여, 에이전트가 학습하지 못한 낯선 환경에서도 강건하게 임무를 수행할 수 있도록 한다. 이 방법은 불확실성이 높은 상태에서 임무 수행에 유리한 공통된 행동을 학습하도록 유도함으로써, 낯선 환경에서 강건한 행동을 가능하게 한다. 4장과 5장에서는 학습 분포 외 상황에 처한 에이전트를 사람의 개입 없이 학습 분포 내 상태로 복귀시키기 위한 재학습 방법을 제안한다. 4장에서는 불확실성 기반 보상을 활용한 자기지도 학습 기법을 소개하고, 5장에서는 대형 언어 모델을 활용하여 학습 분포 외 상황을 분석하고 필요한 복귀 행동을 추론하며, 이에 상응하는 보상을 생성함으로써 에이전트의 복귀를 유도하는 방법을 제시한다.

    이와 같은 방법론들은 강화학습 에이전트의 복원력을 실질적으로 향상시키며, 향후 물리 인공지능 시스템의 발전을 위한 기반을 마련한다.
    번역하기

    본 논문은 강화학습 에이전트의 학습 분포 외 상황에서의 복원력 향상 문제를 다룬다. 심층 강화학습은 다양한 분야에서 뛰어난 성과를 거두었지만, 학습 분포를 벗어난 상황에서는 신뢰할 ...

    본 논문은 강화학습 에이전트의 학습 분포 외 상황에서의 복원력 향상 문제를 다룬다. 심층 강화학습은 다양한 분야에서 뛰어난 성과를 거두었지만, 학습 분포를 벗어난 상황에서는 신뢰할 수 없는 행동을 생성해 임무 수행의 실패로 이어질 수 있다. 기존 연구들은 주로 학습 분포를 확장하거나 에이전트가 학습 분포 외 상황에 처하지 않도록 예방하는 데 초점을 맞추어 왔으나, 실제로 학습 분포 외 상황에 직면했을 때 에이전트가 어떻게 대응해야 하는지에 대해서는 충분히 다루지 않았다.

    이를 해결하기 위해 본 논문에서는 학습 분포 외 상황에서 강화학습 에이전트의 복원력을 정의하고, 이를 향상시키기 위한 새로운 방법론들을 제안한다. 3장에서는 불확실성 기반 조건부 정책을 도입하여, 에이전트가 학습하지 못한 낯선 환경에서도 강건하게 임무를 수행할 수 있도록 한다. 이 방법은 불확실성이 높은 상태에서 임무 수행에 유리한 공통된 행동을 학습하도록 유도함으로써, 낯선 환경에서 강건한 행동을 가능하게 한다. 4장과 5장에서는 학습 분포 외 상황에 처한 에이전트를 사람의 개입 없이 학습 분포 내 상태로 복귀시키기 위한 재학습 방법을 제안한다. 4장에서는 불확실성 기반 보상을 활용한 자기지도 학습 기법을 소개하고, 5장에서는 대형 언어 모델을 활용하여 학습 분포 외 상황을 분석하고 필요한 복귀 행동을 추론하며, 이에 상응하는 보상을 생성함으로써 에이전트의 복귀를 유도하는 방법을 제시한다.

    이와 같은 방법론들은 강화학습 에이전트의 복원력을 실질적으로 향상시키며, 향후 물리 인공지능 시스템의 발전을 위한 기반을 마련한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This dissertation addresses the challenge of enhancing the resilience of reinforcement learning (RL) agents in out-of-distribution (OOD) situations. While deep reinforcement learning (DRL) has achieved notable success across various domains, it struggles to generate reliable actions in OOD scenarios, leading to potential task failures. Previous approaches have primarily focused on expanding the learning distribution or preventing the agent from encountering OOD situations. However, they do not adequately address how the agent should respond when faced with such situations.

    To tackle this gap, the dissertation defines the resilience of RL agents in OOD situations and proposes novel methodologies to enhance it. Chapter 3 introduces an uncertainty-conditioned policy that enables agents to perform tasks robustly in OOD situations. This approach allows agents to learn common actions for high-uncertainty states, ensuring task robustness even in unfamiliar environments. Chapters 4 and 5 propose methods to retrain the agent and guide it back to a state within the learning distribution without human intervention when encountering OOD situations. Specifically, Chapter 4 introduces a self-supervised learning approach using an uncertainty-based auxiliary reward, while Chapter 5 presents a method that leverages large language models (LLMs) to analyze OOD situations, infer necessary recovery behaviors, and generate corresponding rewards to guide the agent's recovery.

    Together, this dissertation contributes to the advancement of more resilient RL agents, laying a foundation for future research in physical artificial intelligence systems.
    번역하기

    This dissertation addresses the challenge of enhancing the resilience of reinforcement learning (RL) agents in out-of-distribution (OOD) situations. While deep reinforcement learning (DRL) has achieved notable success across various domains, it strugg...

    This dissertation addresses the challenge of enhancing the resilience of reinforcement learning (RL) agents in out-of-distribution (OOD) situations. While deep reinforcement learning (DRL) has achieved notable success across various domains, it struggles to generate reliable actions in OOD scenarios, leading to potential task failures. Previous approaches have primarily focused on expanding the learning distribution or preventing the agent from encountering OOD situations. However, they do not adequately address how the agent should respond when faced with such situations.

    To tackle this gap, the dissertation defines the resilience of RL agents in OOD situations and proposes novel methodologies to enhance it. Chapter 3 introduces an uncertainty-conditioned policy that enables agents to perform tasks robustly in OOD situations. This approach allows agents to learn common actions for high-uncertainty states, ensuring task robustness even in unfamiliar environments. Chapters 4 and 5 propose methods to retrain the agent and guide it back to a state within the learning distribution without human intervention when encountering OOD situations. Specifically, Chapter 4 introduces a self-supervised learning approach using an uncertainty-based auxiliary reward, while Chapter 5 presents a method that leverages large language models (LLMs) to analyze OOD situations, infer necessary recovery behaviors, and generate corresponding rewards to guide the agent's recovery.

    Together, this dissertation contributes to the advancement of more resilient RL agents, laying a foundation for future research in physical artificial intelligence systems.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1.1 Background and Motivations 1
    • 1.2 Contributions 6
    • 1.2.1 Enhancing Robustness in Out-of-Distribution 6
    • 1.2.2 Enabling Recovery from Out-of-Distribution through Retraining 8
    • 1 Introduction 1
    • 1.1 Background and Motivations 1
    • 1.2 Contributions 6
    • 1.2.1 Enhancing Robustness in Out-of-Distribution 6
    • 1.2.2 Enabling Recovery from Out-of-Distribution through Retraining 8
    • 2 Preliminaries 11
    • 2.1 Out-of-Distribution in Deep Learning 11
    • 2.2 Epistemic Uncertainty Estimation 12
    • 2.3 Markov Decision Process 14
    • 2.4 Soft Actor-Critic 15
    • 2.5 Mutual Information 16
    • 3 Enhancing Agent Robustness in Out-of-Distribution 17
    • 3.1 Introduction 17
    • 3.2 Related Work 21
    • 3.2.1 Reinforcement Learning for Continuous Control 21
    • 3.2.2 Mutual Information for Structured Behavior 22
    • 3.2.3 Epistemic Uncertainty in Reinforcement Learning 22
    • 3.3 Method 23
    • 3.3.1 Uncertainty Mapping Function 23
    • 3.3.2 Mutual Information for Conditioning Uncertainty 25
    • 3.3.3 Uncertainty-Conditioned Policy 26
    • 3.4 Experiments 29
    • 3.4.1 Environments and Scenarios Details 29
    • 3.4.2 Evaluation in MuJoCo 31
    • 3.4.3 Evaluation in CARLA 35
    • 3.5 Conclusion 36
    • 4 Enabling Agent Recovery from Out-of-Distribution through Self-Supervised Retraining 37
    • 4.1 Introduction 37
    • 4.2 Related Work 39
    • 4.2.1 Preventing OOD Situations in RL 39
    • 4.2.2 Learning to Return to Target State Distributions in RL 40
    • 4.3 Method 42
    • 4.3.1 Auxiliary Reward for Recovery From OOD Situations 43
    • 4.3.2 Uncertainty-Aware Policy Consolidation (UPC) 44
    • 4.3.3 Self-Supervised RL for Recovery From OOD Situations 45
    • 4.4 Experiments 47
    • 4.4.1 Environments 47
    • 4.4.2 Analysis of the Necessity of Retraining 51
    • 4.4.3 Analysis of the Uncertainty Distance 53
    • 4.4.4 Retraining for Recovery From OOD Situations 54
    • 4.4.5 Retraining for Recovery From OOD Situations With the Agent’s Own Criterion 58
    • 4.4.6 Analysis of Each Component of SeRO 61
    • 4.5 Conclusion 63
    • 5 Facilitating Agent Recovery from Out-of-Distribution using Language Models 64
    • 5.1 Introduction 64
    • 5.2 Related Work 67
    • 5.2.1 Addressing OOD in RL 67
    • 5.2.2 Language Models for Reward Generation 68
    • 5.3 Method 70
    • 5.3.1 Prerequisites 70
    • 5.3.2 Recovery Reward Generation 70
    • 5.3.3 Language Model-Guided Policy Consolidation 72
    • 5.4 Experiments 75
    • 5.4.1 Environmental Details 76
    • 5.4.2 Evaluation in Four MuJoCo Environments 77
    • 5.4.3 Evaluation in Complex Environments 79
    • 5.4.4 Qualitative Analysis of Recovery Reward Generation 82
    • 5.4.5 Ablation Study 86
    • 5.5 Prompts 87
    • 5.5.1 Environment Description 87
    • 5.5.2 Full Prompts 87
    • 5.6 Conclusion 88
    • 6 Conclusion 94
    • Abstract (In Korean) 110
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼