Multi-agent reinforcement learning (MARL) is a key technique for cooperative autonomous systems, but it remains difficult when agents must act from incomplete local observations while coordinating in a shared continuous-control task. In such setti...
Multi-agent reinforcement learning (MARL) is a key technique for cooperative autonomous systems, but it remains difficult when agents must act from incomplete local observations while coordinating in a shared continuous-control task. In such settings, each agent observes only part of the environment, the apparent dynamics are non-stationary because other agents are learning at the same time, and simple unimodal action policies may be insufficient for coordinated behavior. This thesis studies these challenges in a field-of-view (FOV)-limited multi-agent control setting and proposes a Multi-Agent Transformer-Diffusion-PPO (MATDP) reinforcement learning framework.
The proposed framework formulates the task as a decentralized partially observable Markov decision process. Each actor receives a short history of local observations masked by a FOV radius, while a centralized critic uses unmasked observations during training. The actor first encodes each agent's observation history with a Transformer encoder and then applies a team-level Transformer to produce a cross-agent latent representation. Conditioned on this representation, a joint diffusion policy generates the 15-dimensional action vector of all three agents together. The policy is optimized online with Proximal Policy Optimization (PPO), using step-average diffusion log probabilities, a centralized value function, potential-based reward shaping, and an assignment auxiliary objective for training stabilization.
Experiments are conducted on the MPE simple_spread benchmark with three agents and three landmarks. The main evaluation uses a four-stage flow that gradually reduces the actor FOV radius from 2.0 to 1.5, 1.0, and 0.5. The proposed MATDP model is compared with MATD, which removes online PPO updates, and MATG, which replaces the diffusion policy with a Gaussian policy, under the same training data and flow setting. In the main experiments, MATDP achieves the highest mean success rate over the final 100 episodes at every FOV stage, obtaining 0.190, 0.170, 0.150, and 0.150, respectively. MATD performs substantially worse in the hardest stage, showing that online policy adaptation is necessary.
The results also identify a clear limitation. In the extreme FOV=0.5 condition, MATDP can still produce stochastic successful episodes, but deterministic evaluation remains weak. Additional diagnostics show that actor-side memory and stronger exploration provide only limited improvement, suggesting that the main bottleneck is insufficient information flow under severe partial observability. Therefore, this thesis concludes that Transformer-Diffusion policies are promising for partially observable multi-agent control, while robust deterministic coordination in extreme visibility-limited settings requires future work on communication, stronger memory, and richer evaluation environments.
Keywords: multi-agent reinforcement learning, transformer, diffusion policy, proximal policy optimization, partial observability, centralized training decentralized execution