본 연구에서는 험지 자율주행을 위해, 대규모 데이터로 사전 학습된 깊이 추정 네트워크를 활용하여 단안 카메라 이미지로부터 BEV(Bird’s-Eye View) 주행가능영역과 지형 높이 맵을 동시에 예...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17372616
서울 : 국민대학교 자동차모빌리티대학원, 2025
학위논문(석사) -- 국민대학교 자동차모빌리티대학원 , 자동차IT융합전공 , 2026. 2
2025
한국어
자율주행 ; 주행가능영역 ; 딥러닝 ; 컴퓨터 비전 ; Autonomous Driving ; Drivable Area ; Deep Learning ; Computer Vision
서울
iv, 60 ; 26 cm
지도교수: 임세준
I804:11014-200000951584
0
상세조회0
다운로드본 연구에서는 험지 자율주행을 위해, 대규모 데이터로 사전 학습된 깊이 추정 네트워크를 활용하여 단안 카메라 이미지로부터 BEV(Bird’s-Eye View) 주행가능영역과 지형 높이 맵을 동시에 예...
본 연구에서는 험지 자율주행을 위해, 대규모 데이터로 사전 학습된 깊이 추정 네트워크를 활용하여 단안 카메라 이미지로부터 BEV(Bird’s-Eye View) 주행가능영역과 지형 높이 맵을 동시에 예측하는 심층 신경망 모델을 제안한다. 기존 도심 환경 중심의 연구는 도로 경계가 불분명한 험지 적용이 어려우며, 라이다(LiDAR) 센서는 비용 및 내구성 문제가 있다.
제안하는 모델은 ResNet-50의 의미론적 정보와 Depth Pro의 3D 기하학적 정보를 BEV 공간으로 투영 및 융합함으로써 입력 특징의 표현력을 강화한다. 이후 예측된 연속적인 지형 높이 맵을 중간 단계에서 주행가능영역 예측 헤드의 입력으로 다시 연결하는 구조를 통해, 지형의 기하학적 구조를 고려한 강건한 주행가능영역 탐지를 수행한다.
RELLIS-3D 험지 주행 데이터셋을 활용한 실험 결과, 제안 모델은 기존 방식 대비 높은 F1-Score를 달성했다. 특히 도심 환경 대비 데이터가 부족한 험지 상황에서, 대규모 사전 학습된 지식을 전이하여 활용하는 전략이 단안 카메라만으로도 안정적인 성능을 확보하는 데 유효함을 확인했다. 또한, 예측된 지형 높이 맵은 실제 LiDAR 정답 데이터와 낮은 오차를 보여 비용 맵(Costmap)으로의 활용 가능성을 확인하였다. 본 연구는 추론 시 고가의 LiDAR 센서 의존성을 제거하면서도, 험지 환경의 강건한 자율주행에 필수적인 지형 정보와 주행 영역을 동시에 제공하는 실용적인 인식 시스템을 제시한다.
다국어 초록 (Multilingual Abstract)
This study proposes a deep neural network model that simultaneously predicts the Bird’s-Eye View (BEV) drivable area and terrain height map using a monocular camera image. This approach aims to overcome the high cost of LiDAR sensors and the limitat...
This study proposes a deep neural network model that simultaneously predicts the Bird’s-Eye View (BEV) drivable area and terrain height map using a monocular camera image. This approach aims to overcome the high cost of LiDAR sensors and the limitations of off-road environments where road boundaries are unclear.
The proposed model utilizes a pre-trained depth estimation network to enhance performance. Specifically, it projects and fuses semantic information from ResNet-50 and 3D geometric information from Depth Pro into the BEV space to strengthen the input features. Afterward, a 'BEV Encoder' learns the surrounding context and predicts a continuous terrain height map as an intermediate step. This predicted height information is then connected to the input of the 'drivable area prediction head,' enabling robust detection that considers the geometric structure of the terrain.
The performance of the model was validated using the RELLIS-3D off-road dataset. Experimental results showed that the proposed model achieved a high F1-Score compared to existing methods. In particular, the strategy of transferring large-scale pre-trained knowledge proved effective in securing stable performance even with limited off-road data. Additionally, the predicted terrain height map demonstrated low error against ground truth LiDAR data, verifying its potential as a costmap. This study presents a practical perception system that eliminates the dependency on expensive LiDAR sensors during inference while providing essential terrain information and drivable areas for robust autonomous driving.
목차 (Table of Contents)