본 논문은 강화학습 에이전트의 학습 분포 외 상황에서의 복원력 향상 문제를 다룬다. 심층 강화학습은 다양한 분야에서 뛰어난 성과를 거두었지만, 학습 분포를 벗어난 상황에서는 신뢰할 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 논문은 강화학습 에이전트의 학습 분포 외 상황에서의 복원력 향상 문제를 다룬다. 심층 강화학습은 다양한 분야에서 뛰어난 성과를 거두었지만, 학습 분포를 벗어난 상황에서는 신뢰할 ...
본 논문은 강화학습 에이전트의 학습 분포 외 상황에서의 복원력 향상 문제를 다룬다. 심층 강화학습은 다양한 분야에서 뛰어난 성과를 거두었지만, 학습 분포를 벗어난 상황에서는 신뢰할 수 없는 행동을 생성해 임무 수행의 실패로 이어질 수 있다. 기존 연구들은 주로 학습 분포를 확장하거나 에이전트가 학습 분포 외 상황에 처하지 않도록 예방하는 데 초점을 맞추어 왔으나, 실제로 학습 분포 외 상황에 직면했을 때 에이전트가 어떻게 대응해야 하는지에 대해서는 충분히 다루지 않았다.
이를 해결하기 위해 본 논문에서는 학습 분포 외 상황에서 강화학습 에이전트의 복원력을 정의하고, 이를 향상시키기 위한 새로운 방법론들을 제안한다. 3장에서는 불확실성 기반 조건부 정책을 도입하여, 에이전트가 학습하지 못한 낯선 환경에서도 강건하게 임무를 수행할 수 있도록 한다. 이 방법은 불확실성이 높은 상태에서 임무 수행에 유리한 공통된 행동을 학습하도록 유도함으로써, 낯선 환경에서 강건한 행동을 가능하게 한다. 4장과 5장에서는 학습 분포 외 상황에 처한 에이전트를 사람의 개입 없이 학습 분포 내 상태로 복귀시키기 위한 재학습 방법을 제안한다. 4장에서는 불확실성 기반 보상을 활용한 자기지도 학습 기법을 소개하고, 5장에서는 대형 언어 모델을 활용하여 학습 분포 외 상황을 분석하고 필요한 복귀 행동을 추론하며, 이에 상응하는 보상을 생성함으로써 에이전트의 복귀를 유도하는 방법을 제시한다.
이와 같은 방법론들은 강화학습 에이전트의 복원력을 실질적으로 향상시키며, 향후 물리 인공지능 시스템의 발전을 위한 기반을 마련한다.
다국어 초록 (Multilingual Abstract)
This dissertation addresses the challenge of enhancing the resilience of reinforcement learning (RL) agents in out-of-distribution (OOD) situations. While deep reinforcement learning (DRL) has achieved notable success across various domains, it strugg...
This dissertation addresses the challenge of enhancing the resilience of reinforcement learning (RL) agents in out-of-distribution (OOD) situations. While deep reinforcement learning (DRL) has achieved notable success across various domains, it struggles to generate reliable actions in OOD scenarios, leading to potential task failures. Previous approaches have primarily focused on expanding the learning distribution or preventing the agent from encountering OOD situations. However, they do not adequately address how the agent should respond when faced with such situations.
To tackle this gap, the dissertation defines the resilience of RL agents in OOD situations and proposes novel methodologies to enhance it. Chapter 3 introduces an uncertainty-conditioned policy that enables agents to perform tasks robustly in OOD situations. This approach allows agents to learn common actions for high-uncertainty states, ensuring task robustness even in unfamiliar environments. Chapters 4 and 5 propose methods to retrain the agent and guide it back to a state within the learning distribution without human intervention when encountering OOD situations. Specifically, Chapter 4 introduces a self-supervised learning approach using an uncertainty-based auxiliary reward, while Chapter 5 presents a method that leverages large language models (LLMs) to analyze OOD situations, infer necessary recovery behaviors, and generate corresponding rewards to guide the agent's recovery.
Together, this dissertation contributes to the advancement of more resilient RL agents, laying a foundation for future research in physical artificial intelligence systems.
목차 (Table of Contents)