The surge of space objects in Earth orbit has intensified collision risks and heightened the strategic importance of on-orbit proximity operations. In this context, the orbital pursuit–evasion game (OPEG) provides a framework for analyzing non-coope...
The surge of space objects in Earth orbit has intensified collision risks and heightened the strategic importance of on-orbit proximity operations. In this context, the orbital pursuit–evasion game (OPEG) provides a framework for analyzing non-cooperative rendezvous scenarios in which one spacecraft attempts to approach another while the target seeks to evade. This thesis presents a RL-based evasive guidance methodology for spacecraft operating in J2-perturbed elliptic orbits.
The proposed framework formulates the evasive guidance problem based on the theoretical foundations of OPEG as a two-player differential game. The evader learns a reward-optimal evasion policy using the Soft Actor-Critic (SAC) algorithm—an entropy-regularized off-policy method that balances exploration and exploitation and promotes robust policy convergence. In the simulation environment for training and evaluation, the pursuer employs a time-varying linear quadratic regulator (TVLQR) within the relative dynamics framework, based on the Gim–Alfriend state transition matrix (GA-STM), which captures the effects of J2 secular drift and orbital eccentricity on the propagation of relative motion. The pursuer's TVLQR controller is designed to minimize a quadratic cost function that penalizes both state deviation and control effort, thereby generating quadratic-effort near-optimal interception trajectories. The evader's reward function is structured to maximize terminal separation while penalizing excessive propellant consumption, promoting fuel-efficient evasion strategies.
Training was conducted over 20 million timesteps under the following conditions: the evader's maximum Delta-v capability was set to 0.15 m/s per decision step (60% of the pursuer's 0.25 m/s), with a total Delta-v budget of 150 m/s, and domain randomization was applied across diverse initial orbital configurations. Under these conditions, the learned policy achieved an evasion success rate of 96% in Monte Carlo evaluations comprising 100 independent scenarios with randomized pursuer initial states. Notably, the evader's mean Delta-v expenditure was approximately 40 m/s—roughly 40% of the pursuer's mean consumption of 110 m/s—demonstrating that the learned policy effectively exploits orbital dynamics to achieve asymmetric fuel efficiency despite possessing inferior maneuvering capability.
To assess operational applicability, the trained policy was applied to a conjunction scenario derived from an Active Debris Removal (ADR) mission targeting KOMPSAT-1, using two-line element (TLE)-based orbital data. The RL-based guidance increased the global minimum distance at closest approach (DCA) from 7.85 km to 9.66 km—an improvement of approximately 23%—thereby demonstrating the potential of the proposed methodology for enhancing collision avoidance performance in realistic mission contexts.