LiDAR 기반 3차원 객체 탐지에서 사용되는 범위 이미지(Range View image)는 물체의 이미지 상의 크기 변화(Scale variation), 인접한 물체 간의 가림(Occlusion)문제를 갖고 있다. 기존에는 네트워크 설계...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17503665
천안 : 한국기술교육대학교 일반대학원, 2026
학위논문(석사) -- 한국기술교육대학교 일반대학원 , 컴퓨터공학과 컴퓨터공학전공 , 2026. 2
2026
한국어
충청남도
; 26 cm
지도교수: 김덕수
I804:44013-200000971177
0
상세조회0
다운로드LiDAR 기반 3차원 객체 탐지에서 사용되는 범위 이미지(Range View image)는 물체의 이미지 상의 크기 변화(Scale variation), 인접한 물체 간의 가림(Occlusion)문제를 갖고 있다. 기존에는 네트워크 설계...
LiDAR 기반 3차원 객체 탐지에서 사용되는 범위 이미지(Range View image)는 물체의 이미지 상의 크기 변화(Scale variation), 인접한 물체 간의 가림(Occlusion)문제를 갖고 있다. 기존에는 네트워크 설계나 범위 이미지에서 변형하는 방법들을 통해 부분적으로 완화하였다. 하지만 범위 이미지라는 입력에 의존함에 따라 이러한 문제를 해결하는데 제약이 존재한다.
본 연구에서는 이러한 문제를 해결하기 위해 다중 시점 직교 투영을 기반으로 한 새로운 입력 이미지와 네트워크 설계를 제안한다. ORV는 관측 방향과 직교하는 평면으로 포인트 클라우드를 투영하여 물체의 형상을 보존하며, 여러 각도에서 ORV(Orthogonal Range View)를 생성하여 단일 직교 투영에서 발생하는 포인트 손실을 크게 줄인다. 또한 ORV의 높은 희소성과 불균일한 이웃 관계를 해결하기 위해, 합성곱(convolution) 연산 시 참조되는 이웃 픽셀을 3차원 거리 기반으로 동적으로 선택하는 MetaDRB(Meta Dilation Residual Block)을 설계하였다. 또한 학습 기반 가중 평균을 이용한 경량 다중 시점 특징 집계 모듈을 제안하여 다중 시점 정보를 효과적으로 통합한다.
KITTI 3D 객체 탐지 실험 결과, 제안하는 ORVrcnn은 RangeRCNN대비 Cyclist에서 최대 +5.1%의 향상, Car 클래스에서도 RangeRCNN과 유사한 수준의 정확도를 유지하면서 전반적인 평균 성능이 개선되었다. 또한, RangeRCNN과의 거리 별 성능 비교에서도 근거리에서는 RangeRCNN과 유사한 수준의 정확도를 보이고, 중거리 구간에서는 일관되게 RangeRCNN대비 성능향상이 있는것을 확인하였다. 추가적으로 ACDet의 범위 이미지 입력과 백본 부분을 본 연구에서 제안한 구조와 입력을 다중 시점 직교 투영으로 대체한 실험에서도 기존 대비 성능이 상승하여, 제안한 ORV 표현의 일반성과 확장성을 확인하였다.
본 연구는 범위 이미지 기반 3D 객체 탐지의 구조적 문제를 완화하는 새로운 투영 방식과 백본(Backbone) 구조를 제시하며, 다양한 범위 이미지 계열 모델에서 적용 가능한 효과적인 투영 대안임을 실험적으로 검증하였다.
다국어 초록 (Multilingual Abstract)
Range-image–based LiDAR 3D object detection suffers from structural limitations such as scale variation in image space and occlusion between adjacent objects. Prior studies have attempted to alleviate these issues through network design or input tra...
Range-image–based LiDAR 3D object detection suffers from structural limitations such as scale variation in image space and occlusion between adjacent objects. Prior studies have attempted to alleviate these issues through network design or input transformations, but fundamental constraints remain as long as the detector relies on the Range image representation.
To address these limitations, this study proposes a new input representation and network architecture based on Multi-view Orthogonal Projection. The proposed Orthogonal View projects the point cloud onto planes orthogonal to the observation direction, preserving the geometric shape of objects. Multiple Orthogonal View projects are generated from different viewing angles to significantly reduce point loss inherent in single orthogonal projection. To overcome the high sparsity and inconsistent neighborhood structure of ORV, we design the Meta Dilation Residual Block (MetaDRB), which dynamically selects dilation neighbors based on 3D distance. Furthermore, we introduce a lightweight Multi-view Feature Aggregation module that integrates multi-view information through a learnable weighted averaging scheme.
Experiments on the KITTI 3D object detection benchmark show that the proposed ORVrcnn improves Cyclist detection performance by up to +5.1\% AP compared to RangeRCNN, while achieving comparable accuracy for Car and improving overall average performance. Distance-wise evaluation further confirms that ORVrcnn maintains similar accuracy in the near range and consistently outperforms RangeRCNN in the mid-range (10–50 m). Additionally, replacing the Range image input and backbone of ACDet with the proposed MOP-based ORV and MetaDRB also leads to performance gains, demonstrating the generality and extensibility of the proposed framework.
Overall, this study introduces a novel projection strategy and backbone design that mitigate key structural issues of Range-image–based 3D object detection, providing an effective and widely applicable alternative representation for Range-view detectors.
목차 (Table of Contents)