RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Transformer-Diffusion-PPO를 이용한 다중 에이전트 강화 학습 연구 = A Study on Multi-Agent Reinforcement Learning using Transformer-Diffusion-PPO

    한글로보기

    https://www.riss.kr/link?id=T17545894

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Multi-agent reinforcement learning (MARL) is a key technique for cooperative autonomous systems, but it remains difficult when agents must act from incomplete local observations while coordinating in a shared continuous-control task. In such settings, each agent observes only part of the environment, the apparent dynamics are non-stationary because other agents are learning at the same time, and simple unimodal action policies may be insufficient for coordinated behavior. This thesis studies these challenges in a field-of-view (FOV)-limited multi-agent control setting and proposes a Multi-Agent Transformer-Diffusion-PPO (MATDP) reinforcement learning framework.
    The proposed framework formulates the task as a decentralized partially observable Markov decision process. Each actor receives a short history of local observations masked by a FOV radius, while a centralized critic uses unmasked observations during training. The actor first encodes each agent's observation history with a Transformer encoder and then applies a team-level Transformer to produce a cross-agent latent representation. Conditioned on this representation, a joint diffusion policy generates the 15-dimensional action vector of all three agents together. The policy is optimized online with Proximal Policy Optimization (PPO), using step-average diffusion log probabilities, a centralized value function, potential-based reward shaping, and an assignment auxiliary objective for training stabilization.
    Experiments are conducted on the MPE simple_spread benchmark with three agents and three landmarks. The main evaluation uses a four-stage flow that gradually reduces the actor FOV radius from 2.0 to 1.5, 1.0, and 0.5. The proposed MATDP model is compared with MATD, which removes online PPO updates, and MATG, which replaces the diffusion policy with a Gaussian policy, under the same training data and flow setting. In the main experiments, MATDP achieves the highest mean success rate over the final 100 episodes at every FOV stage, obtaining 0.190, 0.170, 0.150, and 0.150, respectively. MATD performs substantially worse in the hardest stage, showing that online policy adaptation is necessary.
    The results also identify a clear limitation. In the extreme FOV=0.5 condition, MATDP can still produce stochastic successful episodes, but deterministic evaluation remains weak. Additional diagnostics show that actor-side memory and stronger exploration provide only limited improvement, suggesting that the main bottleneck is insufficient information flow under severe partial observability. Therefore, this thesis concludes that Transformer-Diffusion policies are promising for partially observable multi-agent control, while robust deterministic coordination in extreme visibility-limited settings requires future work on communication, stronger memory, and richer evaluation environments.
    Keywords: multi-agent reinforcement learning, transformer, diffusion policy, proximal policy optimization, partial observability, centralized training decentralized execution
    번역하기

    Multi-agent reinforcement learning (MARL) is a key technique for cooperative autonomous systems, but it remains difficult when agents must act from incomplete local observations while coordinating in a shared continuous-control task. In such setti...

    Multi-agent reinforcement learning (MARL) is a key technique for cooperative autonomous systems, but it remains difficult when agents must act from incomplete local observations while coordinating in a shared continuous-control task. In such settings, each agent observes only part of the environment, the apparent dynamics are non-stationary because other agents are learning at the same time, and simple unimodal action policies may be insufficient for coordinated behavior. This thesis studies these challenges in a field-of-view (FOV)-limited multi-agent control setting and proposes a Multi-Agent Transformer-Diffusion-PPO (MATDP) reinforcement learning framework.
    The proposed framework formulates the task as a decentralized partially observable Markov decision process. Each actor receives a short history of local observations masked by a FOV radius, while a centralized critic uses unmasked observations during training. The actor first encodes each agent's observation history with a Transformer encoder and then applies a team-level Transformer to produce a cross-agent latent representation. Conditioned on this representation, a joint diffusion policy generates the 15-dimensional action vector of all three agents together. The policy is optimized online with Proximal Policy Optimization (PPO), using step-average diffusion log probabilities, a centralized value function, potential-based reward shaping, and an assignment auxiliary objective for training stabilization.
    Experiments are conducted on the MPE simple_spread benchmark with three agents and three landmarks. The main evaluation uses a four-stage flow that gradually reduces the actor FOV radius from 2.0 to 1.5, 1.0, and 0.5. The proposed MATDP model is compared with MATD, which removes online PPO updates, and MATG, which replaces the diffusion policy with a Gaussian policy, under the same training data and flow setting. In the main experiments, MATDP achieves the highest mean success rate over the final 100 episodes at every FOV stage, obtaining 0.190, 0.170, 0.150, and 0.150, respectively. MATD performs substantially worse in the hardest stage, showing that online policy adaptation is necessary.
    The results also identify a clear limitation. In the extreme FOV=0.5 condition, MATDP can still produce stochastic successful episodes, but deterministic evaluation remains weak. Additional diagnostics show that actor-side memory and stronger exploration provide only limited improvement, suggesting that the main bottleneck is insufficient information flow under severe partial observability. Therefore, this thesis concludes that Transformer-Diffusion policies are promising for partially observable multi-agent control, while robust deterministic coordination in extreme visibility-limited settings requires future work on communication, stronger memory, and richer evaluation environments.
    Keywords: multi-agent reinforcement learning, transformer, diffusion policy, proximal policy optimization, partial observability, centralized training decentralized execution

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    다중 에이전트 강화 학습은 협력형 자율 시스템을 구현하기 위한 핵심 기술이 지만, 각 에이전트가 제한된 국소 관측만을 이용하여 연속 제어 행동을 결정해야 하는 경우 안정적인 학습이 어렵다. 본 논문은 시야각(Field-of-View, FOV)이 제한된 다중 에이전트 협력 제어 문제를 대상으로, 관측 이력 모델링과 공동 행동 생성을 결합한 Multi-Agent Transformer-Diffusion-PPO MATDP) 강화학습 프레임워크를 제안한다. 제안한 방법은 문제를 분산 부분 관측 마르코프 결정 과정(Dec-POMDP)으로 정식화한다. 실행 단계의 actor는 FOV 마스크가 적용된 국소 관측 이력만을 입력으로 사용하고, 학습 단계의 critic은 중앙집중식 학습-분산 실행(CTDE) 원칙에 따라 마스크가 적용되지 않은 관측을 사용한다. Actor는 각 에이전트의 관측이력을 Transformer encoder로 표현한 뒤, team-level Transformer를 통해 에이전트 간 문맥 정보를 통합한다. 이후 joint diffusion policy가 세 에이전트의 15차원 공동 행동을 한 번에 생성한다. 정책은 Proximal Policy Optimization(PPO)으로 온라인 업데이트되며, diffusion step별 log probability, 중앙집중식 value function, potential-based reward shaping, assignment auxiliary objective를 함께 사용하여 학습 안정성을 높인다.
    실험은 세 개의 에이전트와 세 개의 landmark로 구성된 MPE simple_spread
    환경에서 수행하였다. 주요 평가는 actor의 FOV 반경을 2.0에서 1.5, 1.0, 0.5 로 점차 줄이는 네 단계 FOV 학습 설정으로 구성하였다. 동일한 학습 규모와 FOV 설정에서 제안한 MATDP, 온라인 PPO 업데이트를 제거한 MATD, 그리고 diffusion policy를 Gaussian policy로 대체한 MATG를 비교하였다. 주요 실험에서 MATDP는 모든 FOV 단계에서 final 100 episodes의 평균 success rate가 가장 높았으며, 각 단계의 성공률은 각각 0.190, 0.170, 0.150, 0.150이었다. MATD는 가장 어려운 FOV=0.5 단계에서 크게 성능이 저하되어, 온라인 정책
    적응의 필요성을 보여준다. MATG와 비교했을 때도 MATDP는 통제된 네 단계 FOV 비교에서 가장 안정적인 success-rate profile을 보였다.
    한편 FOV=0.5 조건은 결정론적 관점에서 완전히 해결되지 않았다. MATDP는 확률적 탐색 과정에서는 성공 episode를 생성할 수 있었지만, deterministic evaluation에서는 여전히 낮은 성능을 보였다. 추가 진단 실험에서도 actor-side memory와 더 강한 exploration은 제한적인 개선만을 제공하였다. 따라서 본 연구는 Transformer-Diffusion 기반 정책이 부분 관측 다중 에이전트 제어에 유망함을 보이는 동시에, 극단적인 관측 제한 상황에서는 명시적 통신, 더 강한 memory mechanism, 그리고 더 풍부한 평가 환경이 필요함을 확인하였다.
    주요어: 다중 에이전트 강화 학습, Transformer, diffusion policy, Proximal Policy Optimization, 부분 관측, 중앙집중식 학습-분산 실행
    번역하기

    다중 에이전트 강화 학습은 협력형 자율 시스템을 구현하기 위한 핵심 기술이 지만, 각 에이전트가 제한된 국소 관측만을 이용하여 연속 제어 행동을 결정해야 하는 경우 안정적인 학습이 ...

    다중 에이전트 강화 학습은 협력형 자율 시스템을 구현하기 위한 핵심 기술이 지만, 각 에이전트가 제한된 국소 관측만을 이용하여 연속 제어 행동을 결정해야 하는 경우 안정적인 학습이 어렵다. 본 논문은 시야각(Field-of-View, FOV)이 제한된 다중 에이전트 협력 제어 문제를 대상으로, 관측 이력 모델링과 공동 행동 생성을 결합한 Multi-Agent Transformer-Diffusion-PPO MATDP) 강화학습 프레임워크를 제안한다. 제안한 방법은 문제를 분산 부분 관측 마르코프 결정 과정(Dec-POMDP)으로 정식화한다. 실행 단계의 actor는 FOV 마스크가 적용된 국소 관측 이력만을 입력으로 사용하고, 학습 단계의 critic은 중앙집중식 학습-분산 실행(CTDE) 원칙에 따라 마스크가 적용되지 않은 관측을 사용한다. Actor는 각 에이전트의 관측이력을 Transformer encoder로 표현한 뒤, team-level Transformer를 통해 에이전트 간 문맥 정보를 통합한다. 이후 joint diffusion policy가 세 에이전트의 15차원 공동 행동을 한 번에 생성한다. 정책은 Proximal Policy Optimization(PPO)으로 온라인 업데이트되며, diffusion step별 log probability, 중앙집중식 value function, potential-based reward shaping, assignment auxiliary objective를 함께 사용하여 학습 안정성을 높인다.
    실험은 세 개의 에이전트와 세 개의 landmark로 구성된 MPE simple_spread
    환경에서 수행하였다. 주요 평가는 actor의 FOV 반경을 2.0에서 1.5, 1.0, 0.5 로 점차 줄이는 네 단계 FOV 학습 설정으로 구성하였다. 동일한 학습 규모와 FOV 설정에서 제안한 MATDP, 온라인 PPO 업데이트를 제거한 MATD, 그리고 diffusion policy를 Gaussian policy로 대체한 MATG를 비교하였다. 주요 실험에서 MATDP는 모든 FOV 단계에서 final 100 episodes의 평균 success rate가 가장 높았으며, 각 단계의 성공률은 각각 0.190, 0.170, 0.150, 0.150이었다. MATD는 가장 어려운 FOV=0.5 단계에서 크게 성능이 저하되어, 온라인 정책
    적응의 필요성을 보여준다. MATG와 비교했을 때도 MATDP는 통제된 네 단계 FOV 비교에서 가장 안정적인 success-rate profile을 보였다.
    한편 FOV=0.5 조건은 결정론적 관점에서 완전히 해결되지 않았다. MATDP는 확률적 탐색 과정에서는 성공 episode를 생성할 수 있었지만, deterministic evaluation에서는 여전히 낮은 성능을 보였다. 추가 진단 실험에서도 actor-side memory와 더 강한 exploration은 제한적인 개선만을 제공하였다. 따라서 본 연구는 Transformer-Diffusion 기반 정책이 부분 관측 다중 에이전트 제어에 유망함을 보이는 동시에, 극단적인 관측 제한 상황에서는 명시적 통신, 더 강한 memory mechanism, 그리고 더 풍부한 평가 환경이 필요함을 확인하였다.
    주요어: 다중 에이전트 강화 학습, Transformer, diffusion policy, Proximal Policy Optimization, 부분 관측, 중앙집중식 학습-분산 실행

    더보기

    목차 (Table of Contents)

    • I. Introduction 1
    • 1. Background and Motivation 1
    • 2. Problem Statement 2
    • 3. Research Objectives 4
    • 4. Main Contributions 4
    • I. Introduction 1
    • 1. Background and Motivation 1
    • 2. Problem Statement 2
    • 3. Research Objectives 4
    • 4. Main Contributions 4
    • 5. Thesis Organization 5
    • II. Related Work 7
    • 1. Deep Reinforcement Learning and Policy Optimization 8
    • 2. Multi-Agent Reinforcement Learning under Partial Observability 9
    • 3. Transformer-Based Sequence Modeling 10
    • 4. Diffusion Policies for Continuous Action Generation 12
    • 5. Reward Shaping and Training Stabilization 13
    • 6. Chapter Summary 14
    • III. Proposed Methodology 15
    • 1. Problem Formulation 15
    • 2. FOV-Limited Observation Model 17
    • 3. Actor Architecture 18
    • 1) Per-Agent History Encoder 18
    • 2) Team-Level Cross-Agent Encoder 20
    • 4. Joint Diffusion Policy 21
    • 5. Centralized Critic 22
    • 6. PPO Update for Diffusion Policy 24
    • 7. Reward Shaping and Assignment Auxiliary Learning 26
    • 1) Potential-Based Reward Shaping 26
    • 2) Assignment Auxiliary Objective 27
    • 8. Curriculum Training Procedure 28
    • 9. Chapter Summary 29
    • IV. Experimental Evaluation and Analysis 31
    • 1. Experimental Environment and Setup 31
    • 2. Compared Methods and Evaluation Metrics 34
    • 3. Main flow Results 36
    • 4. MATDP-MAPPO Fixed-FOV Baseline Comparison 41
    • 5. Ablation Studies 46
    • 1) No-pretraining Baseline at FOV=2.0 46
    • 2) Actor-side Memory Diagnostics 47
    • 6. Extreme POMDP Analysis 48
    • 7. Limitations of MATDP Observed in the Experiments 50
    • 8. Chapter Summary 52
    • V. Conclusion and Future Work 55
    • 1. Conclusion 55
    • 2. Contributions 57
    • 3. Potential Application Scenarios and Practical Use 58
    • 4. Limitations and Future Work 59
    • References 63
    • Abstract 67
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼