RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Designing Robust Autonomous Driving Controller Using Mixed Quality Demonstrations = 혼합 시범을 활용한 견고한 자율주행 제어기 설계

    한글로보기

    https://www.riss.kr/link?id=T17314572

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 학위논문은 자율주행 차량의 배치를 방해하는 여러 가지 문제를 해결하기 위해, 안전 보장, 제한된 데이터로부터의 효율적 학습, 장기 주행 시나리오에 대한 일관된 추론이 가능한 견고한 자율주행 제어기를 설계하는 방법을 제시한다. 기존 자율주행 모방 학습 방법은 주로 전문가의 시연 데이터에만 의존하여 새로운 상황이나 위험 상황에 대한 적응력이 부족하고, 현실 세계의 복잡성에 대한 일반화 능력이 제한된 제어기를 생성하는 문제가 있었다.

    이러한 한계를 극복하기 위해 본 논문은 전문가 시연과 의도적으로 부정적인 운전 예시를 결합한 혼합 품질 시연(mixed-quality demonstrations)을 활용하여 자율주행 정책의 견고성, 효율성 및 안전성을 높이는 통합 프레임워크를 제안한다. 먼저, 본 논문에서는 자율주행에 특화된 마르코프 의사결정 과정(Markov Decision Process) 공식을 재검토하고, 행동 복제(behavior cloning), 적대적 모방 학습(adversarial imitation learning), 신뢰 영역 정책 최적화(trust-region policy optimization)에 대한 엄밀한 수학적 기초를 제공한다. 이를 통해 원칙에 입각한 샘플 재가중(principled sample re-weighting)과 명시적인 안전 지향 비용 모델링(explicit safety-oriented cost modeling)의 중요성을 강조한다.

    이러한 기초를 바탕으로 본 논문은 혼합 생성 적대적 모방 학습(Mixed Generative Adversarial Imitation Learning, MixGAIL)을 소개한다. MixGAIL은 훈련 과정에서 기울기 충돌을 방지하도록 설계된 제약 인식 보상 함수(constraint-aware reward function)를 통해 전문가 시연과 부정적 예시를 결합한다. 실험 결과 MixGAIL은 TORCS 시뮬레이션에서 기존의 GAIL 대비 26\% 더 빠르게 수렴하며, 축소된 물리적 RC 차량 플랫폼에서도 전문가 수준의 성능을 달성함을 입증하였다.

    나아가, 많은 양의 라벨링 데이터에 대한 의존성을 더욱 줄이기 위해 반지도 모방 학습(Semi-Supervised Imitation Learning, SSIL)을 제안하며, 변분 오토인코더(VAE)의 재구성 오차(reconstruction error)를 사용하여 데이터 신뢰도를 평가하고 가중치를 부여한다. SSIL은 라벨링된 데이터가 일부만 있는 경우에도 전문가 수준의 운전 성능을 달성할 정도로 학습 효율성을 크게 향상시키며, RC 차량 주행 벤치마크에서 MixGAIL을 능가하는 성과를 보였다.

    알고리즘적 발전을 보완하기 위해 본 논문은 현대 아이오닉 차량을 이용해 수집한 정상 주행 궤적 42,400개와 비정상 주행 궤적 20,355개로 구성된 R3 주행 데이터셋(R3 Driving Dataset)을 기여한다. 광범위한 기준 평가를 통해 혼합 밀도 네트워크(mixture density networks)는 인식적 불확실성(epistemic uncertainty)을 통해 운전자 유발 이상 상황을 효과적으로 탐지하며, 오토인토더(AE) 재구성 오차는 환경적 위험 요인을 강조하여 더 안전한 주행 정책을 수립할 수 있음을 보여준다.

    마지막으로 본 논문은 보상 및 안전 분별기를 제약된 정책 최적화 프레임워크에 통합한 안전 제약 모방 학습(Safety-Constrained Imitation Learning, SafeIL)을 제안한다. SafeIL은 Safety-Gym, MetaDrive, F1tenth를 포함한 다양한 시뮬레이션 및 실제 로봇 플랫폼에서 안전 위반을 크게 줄이면서도 높은 작업 성능을 유지하였다.

    이러한 연구 결과들은 데이터의 다양성, 원칙적 가중화 전략, 명시적인 비용 모델링 및 실제 비정상 이벤트 데이터를 결합함으로써 자율주행 제어기의 일반화 능력, 효율성 및 안전성을 현저히 향상시킬 수 있음을 입증하며, 신뢰할 수 있는 실제 세계 배치를 위한 실질적인 경로를 제시한다.
    번역하기

    본 학위논문은 자율주행 차량의 배치를 방해하는 여러 가지 문제를 해결하기 위해, 안전 보장, 제한된 데이터로부터의 효율적 학습, 장기 주행 시나리오에 대한 일관된 추론이 가능한 견고...

    본 학위논문은 자율주행 차량의 배치를 방해하는 여러 가지 문제를 해결하기 위해, 안전 보장, 제한된 데이터로부터의 효율적 학습, 장기 주행 시나리오에 대한 일관된 추론이 가능한 견고한 자율주행 제어기를 설계하는 방법을 제시한다. 기존 자율주행 모방 학습 방법은 주로 전문가의 시연 데이터에만 의존하여 새로운 상황이나 위험 상황에 대한 적응력이 부족하고, 현실 세계의 복잡성에 대한 일반화 능력이 제한된 제어기를 생성하는 문제가 있었다.

    이러한 한계를 극복하기 위해 본 논문은 전문가 시연과 의도적으로 부정적인 운전 예시를 결합한 혼합 품질 시연(mixed-quality demonstrations)을 활용하여 자율주행 정책의 견고성, 효율성 및 안전성을 높이는 통합 프레임워크를 제안한다. 먼저, 본 논문에서는 자율주행에 특화된 마르코프 의사결정 과정(Markov Decision Process) 공식을 재검토하고, 행동 복제(behavior cloning), 적대적 모방 학습(adversarial imitation learning), 신뢰 영역 정책 최적화(trust-region policy optimization)에 대한 엄밀한 수학적 기초를 제공한다. 이를 통해 원칙에 입각한 샘플 재가중(principled sample re-weighting)과 명시적인 안전 지향 비용 모델링(explicit safety-oriented cost modeling)의 중요성을 강조한다.

    이러한 기초를 바탕으로 본 논문은 혼합 생성 적대적 모방 학습(Mixed Generative Adversarial Imitation Learning, MixGAIL)을 소개한다. MixGAIL은 훈련 과정에서 기울기 충돌을 방지하도록 설계된 제약 인식 보상 함수(constraint-aware reward function)를 통해 전문가 시연과 부정적 예시를 결합한다. 실험 결과 MixGAIL은 TORCS 시뮬레이션에서 기존의 GAIL 대비 26\% 더 빠르게 수렴하며, 축소된 물리적 RC 차량 플랫폼에서도 전문가 수준의 성능을 달성함을 입증하였다.

    나아가, 많은 양의 라벨링 데이터에 대한 의존성을 더욱 줄이기 위해 반지도 모방 학습(Semi-Supervised Imitation Learning, SSIL)을 제안하며, 변분 오토인코더(VAE)의 재구성 오차(reconstruction error)를 사용하여 데이터 신뢰도를 평가하고 가중치를 부여한다. SSIL은 라벨링된 데이터가 일부만 있는 경우에도 전문가 수준의 운전 성능을 달성할 정도로 학습 효율성을 크게 향상시키며, RC 차량 주행 벤치마크에서 MixGAIL을 능가하는 성과를 보였다.

    알고리즘적 발전을 보완하기 위해 본 논문은 현대 아이오닉 차량을 이용해 수집한 정상 주행 궤적 42,400개와 비정상 주행 궤적 20,355개로 구성된 R3 주행 데이터셋(R3 Driving Dataset)을 기여한다. 광범위한 기준 평가를 통해 혼합 밀도 네트워크(mixture density networks)는 인식적 불확실성(epistemic uncertainty)을 통해 운전자 유발 이상 상황을 효과적으로 탐지하며, 오토인토더(AE) 재구성 오차는 환경적 위험 요인을 강조하여 더 안전한 주행 정책을 수립할 수 있음을 보여준다.

    마지막으로 본 논문은 보상 및 안전 분별기를 제약된 정책 최적화 프레임워크에 통합한 안전 제약 모방 학습(Safety-Constrained Imitation Learning, SafeIL)을 제안한다. SafeIL은 Safety-Gym, MetaDrive, F1tenth를 포함한 다양한 시뮬레이션 및 실제 로봇 플랫폼에서 안전 위반을 크게 줄이면서도 높은 작업 성능을 유지하였다.

    이러한 연구 결과들은 데이터의 다양성, 원칙적 가중화 전략, 명시적인 비용 모델링 및 실제 비정상 이벤트 데이터를 결합함으로써 자율주행 제어기의 일반화 능력, 효율성 및 안전성을 현저히 향상시킬 수 있음을 입증하며, 신뢰할 수 있는 실제 세계 배치를 위한 실질적인 경로를 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This dissertation addresses the challenges of autonomous vehicle deployment by designing robust autonomous driving controllers that effectively handle safety assurance, efficient learning from limited data, and coherent reasoning over extended driving scenarios. Existing imitation learning approaches for autonomous driving predominantly rely on expert-only demonstrations, resulting in controllers that lack adaptability to novel or hazardous situations and exhibit limited generalization to real-world complexities.

    To overcome these limitations, this thesis presents a unified framework utilizing mixed-quality demonstrations—combining expert maneuvers with intentionally negative driving examples—to enhance the robustness, efficiency, and safety of autonomous driving policies. Initially, the dissertation revisits the Markov decision process formulation specific to autonomous driving and provides a rigorous mathematical foundation for behavior cloning, adversarial imitation learning, and trust-region policy optimization. This analysis underscores the significance of principled sample re-weighting and explicit safety-oriented cost modeling.

    Building upon these foundations, the thesis introduces Mixed Generative Adversarial Imitation Learning (MixGAIL), which integrates expert demonstrations with negative examples through a constraint-aware reward function designed to avoid gradient conflicts during training. Empirical results demonstrate that MixGAIL achieves a 26\% faster convergence rate compared to traditional GAIL in TORCS simulations and attains expert-level performance on a scaled physical RC car platform.

    To further reduce reliance on extensively labeled data, this work proposes Semi-Supervised Imitation Learning (SSIL), leveraging variational autoencoder (VAE) reconstruction errors to rank and weight data reliability. SSIL significantly improves learning efficiency, achieving expert-level driving performance with small amount of labeled demonstrations, and surpasses MixGAIL on RC-car traversal benchmarks.

    Complementing these algorithmic advances, this dissertation contributes the R3 Driving Dataset, comprising 42,400 routine and 20,355 abnormal driving trajectories collected using a Hyundai Ioniq vehicle. Extensive baseline evaluations indicate that mixture density networks effectively detect driver-induced anomalies through epistemic uncertainty measures, while autoencoder (AE) reconstruction errors highlight environment-driven hazards, thereby informing safer driving policies.

    Finally, the dissertation proposes Safety-Constrained Imitation Learning (SafeIL), integrating reward and safety discriminators within a constrained policy optimization framework. SafeIL substantially mitigates safety violations across diverse simulation and real-world robotics platforms—including Safety-Gym, MetaDrive, and F1tenth—while preserving high task performance.

    Collectively, these contributions demonstrate that combining data diversity, principled weighting strategies, explicit cost modeling, and real-world abnormal event data significantly enhances the generalization, efficiency, and safety of autonomous driving controllers, charting a viable path toward reliable real-world deployment.
    번역하기

    This dissertation addresses the challenges of autonomous vehicle deployment by designing robust autonomous driving controllers that effectively handle safety assurance, efficient learning from limited data, and coherent reasoning over extended driving...

    This dissertation addresses the challenges of autonomous vehicle deployment by designing robust autonomous driving controllers that effectively handle safety assurance, efficient learning from limited data, and coherent reasoning over extended driving scenarios. Existing imitation learning approaches for autonomous driving predominantly rely on expert-only demonstrations, resulting in controllers that lack adaptability to novel or hazardous situations and exhibit limited generalization to real-world complexities.

    To overcome these limitations, this thesis presents a unified framework utilizing mixed-quality demonstrations—combining expert maneuvers with intentionally negative driving examples—to enhance the robustness, efficiency, and safety of autonomous driving policies. Initially, the dissertation revisits the Markov decision process formulation specific to autonomous driving and provides a rigorous mathematical foundation for behavior cloning, adversarial imitation learning, and trust-region policy optimization. This analysis underscores the significance of principled sample re-weighting and explicit safety-oriented cost modeling.

    Building upon these foundations, the thesis introduces Mixed Generative Adversarial Imitation Learning (MixGAIL), which integrates expert demonstrations with negative examples through a constraint-aware reward function designed to avoid gradient conflicts during training. Empirical results demonstrate that MixGAIL achieves a 26\% faster convergence rate compared to traditional GAIL in TORCS simulations and attains expert-level performance on a scaled physical RC car platform.

    To further reduce reliance on extensively labeled data, this work proposes Semi-Supervised Imitation Learning (SSIL), leveraging variational autoencoder (VAE) reconstruction errors to rank and weight data reliability. SSIL significantly improves learning efficiency, achieving expert-level driving performance with small amount of labeled demonstrations, and surpasses MixGAIL on RC-car traversal benchmarks.

    Complementing these algorithmic advances, this dissertation contributes the R3 Driving Dataset, comprising 42,400 routine and 20,355 abnormal driving trajectories collected using a Hyundai Ioniq vehicle. Extensive baseline evaluations indicate that mixture density networks effectively detect driver-induced anomalies through epistemic uncertainty measures, while autoencoder (AE) reconstruction errors highlight environment-driven hazards, thereby informing safer driving policies.

    Finally, the dissertation proposes Safety-Constrained Imitation Learning (SafeIL), integrating reward and safety discriminators within a constrained policy optimization framework. SafeIL substantially mitigates safety violations across diverse simulation and real-world robotics platforms—including Safety-Gym, MetaDrive, and F1tenth—while preserving high task performance.

    Collectively, these contributions demonstrate that combining data diversity, principled weighting strategies, explicit cost modeling, and real-world abnormal event data significantly enhances the generalization, efficiency, and safety of autonomous driving controllers, charting a viable path toward reliable real-world deployment.

    더보기

    목차 (Table of Contents)

    • 1 INTRODUCTION 1
    • 2 Background 9
    • 2.1 Autonomous Driving: Mathematical Formulation and Challenges 10
    • 2.1.1 MDP Formulation 11
    • 2.1.2 Value Functions and Policy Optimization 11
    • 1 INTRODUCTION 1
    • 2 Background 9
    • 2.1 Autonomous Driving: Mathematical Formulation and Challenges 10
    • 2.1.1 MDP Formulation 11
    • 2.1.2 Value Functions and Policy Optimization 11
    • 2.1.3 Core Challenges 12
    • 2.2 Imitation Learning for Autonomous Driving 13
    • 2.2.1 Problem Statement 13
    • 2.2.2 Behavior Cloning 14
    • 2.2.3 Inverse Reinforcement Learning 14
    • 2.2.4 Challenges Specific to Autonomous Driving 15
    • 2.3 Generative Adversarial Imitation Learning 15
    • 2.3.1 Adversarial Training Details 16
    • 2.3.2 Extensions Relevant to This Thesis 17
    • 2.3.3 Strengths and Limitations 17
    • 2.4 Variational Autoencoders and Mixture Density Networks for Anomaly Detection 18
    • 2.4.1 Variational Autoencoders 18
    • 2.4.2 Mixture Density Networks, Epistemic Uncertainty, and Aleatoric Uncertainty 19
    • 2.5 Constrained Markov Decision Processes 20
    • 2.5.1 Problem Formulation 20
    • 2.5.2 Lagrangian Relaxation 21
    • 2.5.3 Chance Constraints and Risk Measures 22
    • 2.5.4 Dynamic Programming and Policy Search Methods 22
    • 2.6 Related Work 24
    • 2.6.1 Constrained Generative Adversarial Networks 24
    • 2.6.2 Leveraging imperfect data 25
    • 2.6.3 Semi-supervised adversarial imitation learning 25
    • 2.6.4 Existing autonomous driving datasets 26
    • 2.6.5 Safe reinforcement learning and CBFs 27
    • 2.7 Summary 27
    • 3 Mixed Generative Adversarial Imitation Learning 29
    • 3.1 Mixed Demonstration Generative Adversarial Imitation Learning 30
    • 3.1.1 Constraint-reward toy example 36
    • 3.1.2 Utilizing Action Information in CoR 37
    • 3.1.3 Negative Demonstration Exploitation in MixGAIL and IRLF 39
    • 3.2 Experimental Setup 40
    • 3.2.1 Simulation Experiment Setup 40
    • 3.2.2 RC Car Experiment Setup 43
    • 3.3 Results and Analysis 44
    • 3.3.1 TORCS Results 44
    • 3.3.2 RC Car Experiment Results 50
    • 3.4 Conclusion 52
    • 4 Semi-Supervised Imitation Learning 55
    • 4.1 Semi-Supervised Imitation Learning 56
    • 4.1.1 Evaluating Unlabeled Demonstrations 57
    • 4.1.2 Leverage Constraint Reward 60
    • 4.2 Experimental Results 63
    • 4.2.1 Simulation Experiment 63
    • 4.2.2 Leverage-Constrained Reward toy example 64
    • 4.2.3 Real-World RC Car Experiment 80
    • 5 Towards Defensive Autonomous Driving: Collecting and Probing Driving Demonstrations of Mixed Qualities 95
    • 5.1 Data Collection Method 97
    • 5.1.1 Object Detection 97
    • 5.1.2 Object Tracking 100
    • 5.1.3 Prediction Uncertainty in a Mixture Density Network 101
    • 5.1.4 Using Autoencoder Results as a Measurement of Uncertainty 103
    • 5.1.5 Overview of the Collected Data 105
    • 5.2 Experiment Setup 106
    • 5.2.1 Hardware Setting 106
    • 5.2.2 Data Collection Method 106
    • 5.2.3 Anomaly Detection 110
    • 5.3 experiment results 112
    • 6 SafeIL: Safety Constrained Imitation Learning for Autonomous Systems 129
    • 6.1 Preliminaries 130
    • 6.2 Safety constrained imitation learning 131
    • 6.3 Experimental Setup 140
    • 6.3.1 Safety-Gym and Jackal Simulator 142
    • 6.3.2 Autonomous Driving 143
    • 6.3.3 Jackal Platform 144
    • 6.3.4 F1tenth simulator and Racecar Platform 144
    • 6.4 Result 145
    • 6.4.1 Safety-Gym and Jackal Simulator 145
    • 6.4.2 Metadrive simulator 150
    • 6.4.3 CARLA simulator 150
    • 6.4.4 F1tenth simulator 151
    • 6.4.5 Real-world Jackal platform 151
    • 6.4.6 Real-world RC car platform 152
    • 7 Conclusion 163
    • Bibliography 167
    • 초록 179
    • 감사의 글 181
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼