RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    다중 시점 직교 투영을 활용한 LiDAR 기반 3차원 객체 탐지 = Multi-View Orthogonal Projection for LiDAR-Based 3D Object Detection

    한글로보기

    https://www.riss.kr/link?id=T17503665

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    LiDAR 기반 3차원 객체 탐지에서 사용되는 범위 이미지(Range View image)는 물체의 이미지 상의 크기 변화(Scale variation), 인접한 물체 간의 가림(Occlusion)문제를 갖고 있다. 기존에는 네트워크 설계나 범위 이미지에서 변형하는 방법들을 통해 부분적으로 완화하였다. 하지만 범위 이미지라는 입력에 의존함에 따라 이러한 문제를 해결하는데 제약이 존재한다.
    본 연구에서는 이러한 문제를 해결하기 위해 다중 시점 직교 투영을 기반으로 한 새로운 입력 이미지와 네트워크 설계를 제안한다. ORV는 관측 방향과 직교하는 평면으로 포인트 클라우드를 투영하여 물체의 형상을 보존하며, 여러 각도에서 ORV(Orthogonal Range View)를 생성하여 단일 직교 투영에서 발생하는 포인트 손실을 크게 줄인다. 또한 ORV의 높은 희소성과 불균일한 이웃 관계를 해결하기 위해, 합성곱(convolution) 연산 시 참조되는 이웃 픽셀을 3차원 거리 기반으로 동적으로 선택하는 MetaDRB(Meta Dilation Residual Block)을 설계하였다. 또한 학습 기반 가중 평균을 이용한 경량 다중 시점 특징 집계 모듈을 제안하여 다중 시점 정보를 효과적으로 통합한다.

    KITTI 3D 객체 탐지 실험 결과, 제안하는 ORVrcnn은 RangeRCNN대비 Cyclist에서 최대 +5.1%의 향상, Car 클래스에서도 RangeRCNN과 유사한 수준의 정확도를 유지하면서 전반적인 평균 성능이 개선되었다. 또한, RangeRCNN과의 거리 별 성능 비교에서도 근거리에서는 RangeRCNN과 유사한 수준의 정확도를 보이고, 중거리 구간에서는 일관되게 RangeRCNN대비 성능향상이 있는것을 확인하였다. 추가적으로 ACDet의 범위 이미지 입력과 백본 부분을 본 연구에서 제안한 구조와 입력을 다중 시점 직교 투영으로 대체한 실험에서도 기존 대비 성능이 상승하여, 제안한 ORV 표현의 일반성과 확장성을 확인하였다.

    본 연구는 범위 이미지 기반 3D 객체 탐지의 구조적 문제를 완화하는 새로운 투영 방식과 백본(Backbone) 구조를 제시하며, 다양한 범위 이미지 계열 모델에서 적용 가능한 효과적인 투영 대안임을 실험적으로 검증하였다.
    번역하기

    LiDAR 기반 3차원 객체 탐지에서 사용되는 범위 이미지(Range View image)는 물체의 이미지 상의 크기 변화(Scale variation), 인접한 물체 간의 가림(Occlusion)문제를 갖고 있다. 기존에는 네트워크 설계...

    LiDAR 기반 3차원 객체 탐지에서 사용되는 범위 이미지(Range View image)는 물체의 이미지 상의 크기 변화(Scale variation), 인접한 물체 간의 가림(Occlusion)문제를 갖고 있다. 기존에는 네트워크 설계나 범위 이미지에서 변형하는 방법들을 통해 부분적으로 완화하였다. 하지만 범위 이미지라는 입력에 의존함에 따라 이러한 문제를 해결하는데 제약이 존재한다.
    본 연구에서는 이러한 문제를 해결하기 위해 다중 시점 직교 투영을 기반으로 한 새로운 입력 이미지와 네트워크 설계를 제안한다. ORV는 관측 방향과 직교하는 평면으로 포인트 클라우드를 투영하여 물체의 형상을 보존하며, 여러 각도에서 ORV(Orthogonal Range View)를 생성하여 단일 직교 투영에서 발생하는 포인트 손실을 크게 줄인다. 또한 ORV의 높은 희소성과 불균일한 이웃 관계를 해결하기 위해, 합성곱(convolution) 연산 시 참조되는 이웃 픽셀을 3차원 거리 기반으로 동적으로 선택하는 MetaDRB(Meta Dilation Residual Block)을 설계하였다. 또한 학습 기반 가중 평균을 이용한 경량 다중 시점 특징 집계 모듈을 제안하여 다중 시점 정보를 효과적으로 통합한다.

    KITTI 3D 객체 탐지 실험 결과, 제안하는 ORVrcnn은 RangeRCNN대비 Cyclist에서 최대 +5.1%의 향상, Car 클래스에서도 RangeRCNN과 유사한 수준의 정확도를 유지하면서 전반적인 평균 성능이 개선되었다. 또한, RangeRCNN과의 거리 별 성능 비교에서도 근거리에서는 RangeRCNN과 유사한 수준의 정확도를 보이고, 중거리 구간에서는 일관되게 RangeRCNN대비 성능향상이 있는것을 확인하였다. 추가적으로 ACDet의 범위 이미지 입력과 백본 부분을 본 연구에서 제안한 구조와 입력을 다중 시점 직교 투영으로 대체한 실험에서도 기존 대비 성능이 상승하여, 제안한 ORV 표현의 일반성과 확장성을 확인하였다.

    본 연구는 범위 이미지 기반 3D 객체 탐지의 구조적 문제를 완화하는 새로운 투영 방식과 백본(Backbone) 구조를 제시하며, 다양한 범위 이미지 계열 모델에서 적용 가능한 효과적인 투영 대안임을 실험적으로 검증하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Range-image–based LiDAR 3D object detection suffers from structural limitations such as scale variation in image space and occlusion between adjacent objects. Prior studies have attempted to alleviate these issues through network design or input transformations, but fundamental constraints remain as long as the detector relies on the Range image representation.

    To address these limitations, this study proposes a new input representation and network architecture based on Multi-view Orthogonal Projection. The proposed Orthogonal View projects the point cloud onto planes orthogonal to the observation direction, preserving the geometric shape of objects. Multiple Orthogonal View projects are generated from different viewing angles to significantly reduce point loss inherent in single orthogonal projection. To overcome the high sparsity and inconsistent neighborhood structure of ORV, we design the Meta Dilation Residual Block (MetaDRB), which dynamically selects dilation neighbors based on 3D distance. Furthermore, we introduce a lightweight Multi-view Feature Aggregation module that integrates multi-view information through a learnable weighted averaging scheme.

    Experiments on the KITTI 3D object detection benchmark show that the proposed ORVrcnn improves Cyclist detection performance by up to +5.1\% AP compared to RangeRCNN, while achieving comparable accuracy for Car and improving overall average performance. Distance-wise evaluation further confirms that ORVrcnn maintains similar accuracy in the near range and consistently outperforms RangeRCNN in the mid-range (10–50 m). Additionally, replacing the Range image input and backbone of ACDet with the proposed MOP-based ORV and MetaDRB also leads to performance gains, demonstrating the generality and extensibility of the proposed framework.

    Overall, this study introduces a novel projection strategy and backbone design that mitigate key structural issues of Range-image–based 3D object detection, providing an effective and widely applicable alternative representation for Range-view detectors.
    번역하기

    Range-image–based LiDAR 3D object detection suffers from structural limitations such as scale variation in image space and occlusion between adjacent objects. Prior studies have attempted to alleviate these issues through network design or input tra...

    Range-image–based LiDAR 3D object detection suffers from structural limitations such as scale variation in image space and occlusion between adjacent objects. Prior studies have attempted to alleviate these issues through network design or input transformations, but fundamental constraints remain as long as the detector relies on the Range image representation.

    To address these limitations, this study proposes a new input representation and network architecture based on Multi-view Orthogonal Projection. The proposed Orthogonal View projects the point cloud onto planes orthogonal to the observation direction, preserving the geometric shape of objects. Multiple Orthogonal View projects are generated from different viewing angles to significantly reduce point loss inherent in single orthogonal projection. To overcome the high sparsity and inconsistent neighborhood structure of ORV, we design the Meta Dilation Residual Block (MetaDRB), which dynamically selects dilation neighbors based on 3D distance. Furthermore, we introduce a lightweight Multi-view Feature Aggregation module that integrates multi-view information through a learnable weighted averaging scheme.

    Experiments on the KITTI 3D object detection benchmark show that the proposed ORVrcnn improves Cyclist detection performance by up to +5.1\% AP compared to RangeRCNN, while achieving comparable accuracy for Car and improving overall average performance. Distance-wise evaluation further confirms that ORVrcnn maintains similar accuracy in the near range and consistently outperforms RangeRCNN in the mid-range (10–50 m). Additionally, replacing the Range image input and backbone of ACDet with the proposed MOP-based ORV and MetaDRB also leads to performance gains, demonstrating the generality and extensibility of the proposed framework.

    Overall, this study introduces a novel projection strategy and backbone design that mitigate key structural issues of Range-image–based 3D object detection, providing an effective and widely applicable alternative representation for Range-view detectors.

    더보기

    목차 (Table of Contents)

    • 1. 서론 1
    • 1.1. 연구 배경 1
    • 1.2. 연구 개요 4
    • 2. 관련 연구 6
    • 2.1. 3차원 객체 탐지 6
    • 1. 서론 1
    • 1.1. 연구 배경 1
    • 1.2. 연구 개요 4
    • 2. 관련 연구 6
    • 2.1. 3차원 객체 탐지 6
    • 3. 제안하는 방법 10
    • 3.1. 개요 10
    • 3.2. 직교 투영 10
    • 3.3. 다중 시점 직교 투영 13
    • 3.4. 모델 구조 개요 15
    • 3.5. MetaDRB 16
    • 3.6. ORV-Fusion 모듈 19
    • 3.7. 손실 함수 21
    • 4. 결과 및 분석 25
    • 4.1. 데이터 셋과 평가 지표 25
    • 4.2. 구현 상세 25
    • 4.3. 정량적 평가 26
    • 4.3.1. Baseline 재구현 27
    • 4.3.2. 3차원 객체 탐지 정확도 비교 27
    • 4.3.3. 물체 거리에 따른 성능 변화 28
    • 4.3.4. 비 범위 이미지 기반 모델에 대한 확장성 30
    • 4.4. 분석 30
    • 4.4.1. 삭마 연구 31
    • 4.4.2. 다중 시점 활용의 효과 32
    • 4.4.3. ORV 해상도의 영향 33
    • 4.4.4. MetaDRB의 효과 34
    • 4.4.5. ORV-Fusion 대 Cross-view attention 35
    • 4.4.6. 실행 시간 분석 36
    • 4.5. 정성적 평가 37
    • 4.5.1. 정성적 결과 분석 38
    • 5. 결론 48
    • 5.1. 결론 및 한계점 48
    • 5.1.1. 결론 48
    • 5.1.2. 한계점 50
    • 5.2. 향후 연구 51
    • 참고문헌 53
    • Abstract 58
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼