RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    RL-Based Evasive Guidance in J2-Perturbed Elliptic Orbital Pursuit-Evasion Game = J2 섭동 타원궤도상 추격-회피 게임의 강화학습 기반 회피 유도

    한글로보기

    https://www.riss.kr/link?id=T17450884

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The surge of space objects in Earth orbit has intensified collision risks and heightened the strategic importance of on-orbit proximity operations. In this context, the orbital pursuit–evasion game (OPEG) provides a framework for analyzing non-cooperative rendezvous scenarios in which one spacecraft attempts to approach another while the target seeks to evade. This thesis presents a RL-based evasive guidance methodology for spacecraft operating in J2-perturbed elliptic orbits.
    The proposed framework formulates the evasive guidance problem based on the theoretical foundations of OPEG as a two-player differential game. The evader learns a reward-optimal evasion policy using the Soft Actor-Critic (SAC) algorithm—an entropy-regularized off-policy method that balances exploration and exploitation and promotes robust policy convergence. In the simulation environment for training and evaluation, the pursuer employs a time-varying linear quadratic regulator (TVLQR) within the relative dynamics framework, based on the Gim–Alfriend state transition matrix (GA-STM), which captures the effects of J2 secular drift and orbital eccentricity on the propagation of relative motion. The pursuer's TVLQR controller is designed to minimize a quadratic cost function that penalizes both state deviation and control effort, thereby generating quadratic-effort near-optimal interception trajectories. The evader's reward function is structured to maximize terminal separation while penalizing excessive propellant consumption, promoting fuel-efficient evasion strategies.
    Training was conducted over 20 million timesteps under the following conditions: the evader's maximum Delta-v capability was set to 0.15 m/s per decision step (60% of the pursuer's 0.25 m/s), with a total Delta-v budget of 150 m/s, and domain randomization was applied across diverse initial orbital configurations. Under these conditions, the learned policy achieved an evasion success rate of 96% in Monte Carlo evaluations comprising 100 independent scenarios with randomized pursuer initial states. Notably, the evader's mean Delta-v expenditure was approximately 40 m/s—roughly 40% of the pursuer's mean consumption of 110 m/s—demonstrating that the learned policy effectively exploits orbital dynamics to achieve asymmetric fuel efficiency despite possessing inferior maneuvering capability.
    To assess operational applicability, the trained policy was applied to a conjunction scenario derived from an Active Debris Removal (ADR) mission targeting KOMPSAT-1, using two-line element (TLE)-based orbital data. The RL-based guidance increased the global minimum distance at closest approach (DCA) from 7.85 km to 9.66 km—an improvement of approximately 23%—thereby demonstrating the potential of the proposed methodology for enhancing collision avoidance performance in realistic mission contexts.
    번역하기

    The surge of space objects in Earth orbit has intensified collision risks and heightened the strategic importance of on-orbit proximity operations. In this context, the orbital pursuit–evasion game (OPEG) provides a framework for analyzing non-coope...

    The surge of space objects in Earth orbit has intensified collision risks and heightened the strategic importance of on-orbit proximity operations. In this context, the orbital pursuit–evasion game (OPEG) provides a framework for analyzing non-cooperative rendezvous scenarios in which one spacecraft attempts to approach another while the target seeks to evade. This thesis presents a RL-based evasive guidance methodology for spacecraft operating in J2-perturbed elliptic orbits.
    The proposed framework formulates the evasive guidance problem based on the theoretical foundations of OPEG as a two-player differential game. The evader learns a reward-optimal evasion policy using the Soft Actor-Critic (SAC) algorithm—an entropy-regularized off-policy method that balances exploration and exploitation and promotes robust policy convergence. In the simulation environment for training and evaluation, the pursuer employs a time-varying linear quadratic regulator (TVLQR) within the relative dynamics framework, based on the Gim–Alfriend state transition matrix (GA-STM), which captures the effects of J2 secular drift and orbital eccentricity on the propagation of relative motion. The pursuer's TVLQR controller is designed to minimize a quadratic cost function that penalizes both state deviation and control effort, thereby generating quadratic-effort near-optimal interception trajectories. The evader's reward function is structured to maximize terminal separation while penalizing excessive propellant consumption, promoting fuel-efficient evasion strategies.
    Training was conducted over 20 million timesteps under the following conditions: the evader's maximum Delta-v capability was set to 0.15 m/s per decision step (60% of the pursuer's 0.25 m/s), with a total Delta-v budget of 150 m/s, and domain randomization was applied across diverse initial orbital configurations. Under these conditions, the learned policy achieved an evasion success rate of 96% in Monte Carlo evaluations comprising 100 independent scenarios with randomized pursuer initial states. Notably, the evader's mean Delta-v expenditure was approximately 40 m/s—roughly 40% of the pursuer's mean consumption of 110 m/s—demonstrating that the learned policy effectively exploits orbital dynamics to achieve asymmetric fuel efficiency despite possessing inferior maneuvering capability.
    To assess operational applicability, the trained policy was applied to a conjunction scenario derived from an Active Debris Removal (ADR) mission targeting KOMPSAT-1, using two-line element (TLE)-based orbital data. The RL-based guidance increased the global minimum distance at closest approach (DCA) from 7.85 km to 9.66 km—an improvement of approximately 23%—thereby demonstrating the potential of the proposed methodology for enhancing collision avoidance performance in realistic mission contexts.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    지구 궤도상 우주물체의 급격한 증가는 충돌 위험을 심화시키고, 궤도상 근접 운용의 전략적 중요성을 부각시켰다. 이러한 맥락에서 궤도 추격-회피 게임(Orbital Pursuit-Evasion Game, OPEG)은 한 우주선이 다른 우주선에 접근을 시도하고 표적이 이를 회피하고자 하는 비협조적 랑데부 시나리오 분석을 위한 엄밀한 이론적 틀을 제공한다. 본 논문은 J2 섭동이 존재하는 타원궤도에서 운용되는 우주선을 위한 강화학습 기반 회피 유도 기법을 제시한다.
    제안된 프레임워크는 2인 미분게임으로서의 OPEG 이론적 토대를 기반으로 회피 유도 문제를 정식화한다. 회피자는 탐색과 활용의 균형을 유지하면서 강건한 정책 수렴을 촉진하는 엔트로피 정규화 비정책(off-policy) 기법인 Soft Actor-Critic(SAC) 알고리즘을 통해 최적 회피 정책을 학습한다. 학습 및 평가를 위한 시뮬레이션 환경에서, 추격자는 Gim–Alfriend 상태천이행렬(GA-STM) 기반의 상대 동역학 체계 내에서 시변 선형이차조절기(TVLQR)를 활용하여 요격 명령을 생성한다. 이 상대 동역학은 J2 장주기 표류 효과와 궤도 이심률이 상대운동 전파에 미치는 영향을 정밀하게 반영한다. 추격자의 TVLQR 제어기는 상태 편차와 제어 노력을 동시에 벌점화하는 2차 비용함수를 최소화하도록 설계되어 준최적 요격 궤적을 생성한다. 회피자의 보상함수는 종단 분리거리를 극대화하면서 과도한 추진제 소모를 억제하도록 구성되어 연료 효율적인 회피 전략 학습을 유도한다.
    학습은 다음 조건 하에서 2천만 타임스텝에 걸쳐 수행되었다. 회피자의 의사결정 스텝당 최대 Delta-V 능력은 0.15 m/s (추격자 0.25 m/s의 60%), 총 Delta-V 예산은 150 m/s로 설정되었으며, 일정한 분리거리 범위 내에서 다양한 초기 궤도 조건에 대한 도메인 샘플링이 적용되었다. 이러한 조건 하에서, 학습된 정책은 추격자 초기 상태를 임의의 100회의 독립 시나리오로 구성된 Monte-Carlo 평가에서 96%의 회피 성공률을 달성하였다. 특히 회피자의 평균 Delta-V 소모량은 약 40 m/s로, 추격자의 평균 소모량 110 m/s의 약 40%에 불과하였다. 이는 회피자가 열세한 기동 능력을 보유함에도 불구하고, 학습된 정책이 궤도 동역학을 효과적으로 활용하여 비대칭적 연료 효율성을 달성함을 입증한다.
    운용 적용 가능성을 평가하기 위해, 학습된 정책을 KOMPSAT-1을 표적으로 하는 능동적 우주잔해물 제거(ADR) 임무로부터 도출된 조우 시나리오에 적용하였으며, 궤도 데이터는 이행요소(TLE)를 기반으로 산출하였다. 강화학습 기반 유도 적용 결과, 전역 최소 최근접거리(DCA)가 7.85 km에서 9.66 km로 약 23% 증가하였으며, 이를 통해 제안된 기법이 실제 임무 환경에서 충돌회피 성능 향상에 기여할 수 있음을 보였다.
    번역하기

    지구 궤도상 우주물체의 급격한 증가는 충돌 위험을 심화시키고, 궤도상 근접 운용의 전략적 중요성을 부각시켰다. 이러한 맥락에서 궤도 추격-회피 게임(Orbital Pursuit-Evasion Game, OPEG)은 한 우...

    지구 궤도상 우주물체의 급격한 증가는 충돌 위험을 심화시키고, 궤도상 근접 운용의 전략적 중요성을 부각시켰다. 이러한 맥락에서 궤도 추격-회피 게임(Orbital Pursuit-Evasion Game, OPEG)은 한 우주선이 다른 우주선에 접근을 시도하고 표적이 이를 회피하고자 하는 비협조적 랑데부 시나리오 분석을 위한 엄밀한 이론적 틀을 제공한다. 본 논문은 J2 섭동이 존재하는 타원궤도에서 운용되는 우주선을 위한 강화학습 기반 회피 유도 기법을 제시한다.
    제안된 프레임워크는 2인 미분게임으로서의 OPEG 이론적 토대를 기반으로 회피 유도 문제를 정식화한다. 회피자는 탐색과 활용의 균형을 유지하면서 강건한 정책 수렴을 촉진하는 엔트로피 정규화 비정책(off-policy) 기법인 Soft Actor-Critic(SAC) 알고리즘을 통해 최적 회피 정책을 학습한다. 학습 및 평가를 위한 시뮬레이션 환경에서, 추격자는 Gim–Alfriend 상태천이행렬(GA-STM) 기반의 상대 동역학 체계 내에서 시변 선형이차조절기(TVLQR)를 활용하여 요격 명령을 생성한다. 이 상대 동역학은 J2 장주기 표류 효과와 궤도 이심률이 상대운동 전파에 미치는 영향을 정밀하게 반영한다. 추격자의 TVLQR 제어기는 상태 편차와 제어 노력을 동시에 벌점화하는 2차 비용함수를 최소화하도록 설계되어 준최적 요격 궤적을 생성한다. 회피자의 보상함수는 종단 분리거리를 극대화하면서 과도한 추진제 소모를 억제하도록 구성되어 연료 효율적인 회피 전략 학습을 유도한다.
    학습은 다음 조건 하에서 2천만 타임스텝에 걸쳐 수행되었다. 회피자의 의사결정 스텝당 최대 Delta-V 능력은 0.15 m/s (추격자 0.25 m/s의 60%), 총 Delta-V 예산은 150 m/s로 설정되었으며, 일정한 분리거리 범위 내에서 다양한 초기 궤도 조건에 대한 도메인 샘플링이 적용되었다. 이러한 조건 하에서, 학습된 정책은 추격자 초기 상태를 임의의 100회의 독립 시나리오로 구성된 Monte-Carlo 평가에서 96%의 회피 성공률을 달성하였다. 특히 회피자의 평균 Delta-V 소모량은 약 40 m/s로, 추격자의 평균 소모량 110 m/s의 약 40%에 불과하였다. 이는 회피자가 열세한 기동 능력을 보유함에도 불구하고, 학습된 정책이 궤도 동역학을 효과적으로 활용하여 비대칭적 연료 효율성을 달성함을 입증한다.
    운용 적용 가능성을 평가하기 위해, 학습된 정책을 KOMPSAT-1을 표적으로 하는 능동적 우주잔해물 제거(ADR) 임무로부터 도출된 조우 시나리오에 적용하였으며, 궤도 데이터는 이행요소(TLE)를 기반으로 산출하였다. 강화학습 기반 유도 적용 결과, 전역 최소 최근접거리(DCA)가 7.85 km에서 9.66 km로 약 23% 증가하였으며, 이를 통해 제안된 기법이 실제 임무 환경에서 충돌회피 성능 향상에 기여할 수 있음을 보였다.

    더보기

    목차 (Table of Contents)

    • ABSTRACT ii
    • LIST OF FIGURES xii
    • LIST OF TABLES xiii
    • ABBREVIATIONS xiv
    • CHAPTER 1.Introduction 1
    • ABSTRACT ii
    • LIST OF FIGURES xii
    • LIST OF TABLES xiii
    • ABBREVIATIONS xiv
    • CHAPTER 1.Introduction 1
    • 1.1 Background and Motivation 1
    • 1.2 Literature Review 8
    • 1.3 Contribution 15
    • 1.4 Organization 18
    • CHAPTER 2.Preliminaries 19
    • 2.1 Problem Statements 19
    • 2.2 Orbital Pursuit-Evasion Game (OPEG) 23
    • 2.3 Relative Dynamics 31
    • 2.4 Reinforcement Learning 41
    • 2.4.1 Markov Decision Process (MDP) 41
    • 2.4.2 Soft Actor-Critic Algorithm 43
    • CHAPTER 3.Methodology 49
    • 3.1 Pursuit Strategy 49
    • 3.1.1 Time-Varying Linear Quadratic Regulator (TVLQR) 49
    • 3.1.2 Controller Design for Pursuit 54
    • 3.2 Evader RL Reward Design 60
    • 3.3 Training Loop 64
    • CHAPTER 4.Simulations 70
    • 4.1 Simulation Environment 70
    • 4.2 Simulation Result 85
    • 4.2.1 Learning Result 85
    • 4.2.2 Monte-Carlo Test 98
    • 4.3 ADR Mission Design and Application 105
    • 4.3.1 ADR Mission Overview 106
    • 4.3.2 Conjunction Assessment and Application 110
    • 4.3.3 Simulation Results and Analysis 117
    • CHAPTER 5.Conclusions 122
    • 5.1 Research Summary 122
    • 5.2 Future Work 125
    • References 127
    • 국문요약 135
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼