RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Scalable Architecture Methodology for High-Efficiency Multi-Modal Object Detection in Remote Sensing = 원격 탐사 고효율 다중 모달 객체 탐지를 위한 확장 가능 아키텍처 방법론 연구

    한글로보기

    https://www.riss.kr/link?id=T17355630

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    학원 원격 탐사 (Remote Sensing, RS) 객체 탐지는 실제 센싱 환경에서 지속적인 어려움을 겪고 있으며, 그 주요 오류 원인은 네트워크 용량 자체보다는 센서 물리 특성과 취득 기하에 의해 지배된다. 합성 개구 레이다 (Synthetic Aperture Radar, SAR) 영상에서는 반점 노이즈 (Speckle)와 약한 객체 경계로 인해 클러터 기반 오경보 (false alarm) 와 위치 추정의 모호성이 발생한다. 무인항공기 (Unmanned Aerial Vehicle, UAV) 기반 광학 영상에서는 극소형 객체, 높은 공간 밀집도, 그리고 빈번한 가림 현상으로 인해 특징 표현이 심각하게 저하되며 탐지 실패가 증가한다. 또한 다중 센싱 모달리티 (Multi-modality sensing)를 공동으로 활용하는 경우, 서로 다른 영상 형성 메커니즘으로 인해 표현 불일치, 공간 매칭의 불완전함, 그리고 과업 수준의 불일치를 초래할 수 있으며, 단순한 초기 융합 전략은 불안정할 수 있고 심지어 성능 저하를 초래할 수 있다. 이러한 문제를 해결하기 위해 본 학위논문은 원격 탐사 객체 탐지를 위한 확장 가능하고 제약 기반의 고효율 아키텍처 방법론을 제안한다. 제안된 방법론은 탐지 성능을 모델 용량이나 융합 복잡도의 함수로 간주하는 기존 전통적인 방법과 달리, 센싱으로 유발된 불일치가 탐지 프로세스의 어느 단계에서 발생하는지를 명확하게 분석하고, 이를 세 단계를 통해 점진적으로 제거한다. 구체적으로, 모달리티별 표현 안정화, 제어된 경량 교차 모달리티 상호 작용 및 탐지 헤드에서의 과업 일관적 예측을 통해 불일치를 해결한다. 이러한 단계 인지적 (stage-aware) 설계는 단일 모달리티 최적화에서 교차 모달 통합으로 자연스럽게 확장 가능한 통합 설계 논리를 제공하며, 불필요한 계산 오버헤드 (overhead)를 초래하지 않는다. 제안된 방법론은 세 가지 대표적인 객체 탐지 프레임워크를 통해 구체화되고 검증된다. SMEP-DETR은 SAR 영상에 특화된 트랜스포머 기반 (transformer-based) 탐지 모델로서, 다중 에지 (edge) 강화와 병렬 팽창 문맥 집계를 통해 반점 노이즈 안정화와 명시적 경계 강화를 수행함으로써 해양 환경에서의 탐지 정확도를 향상시킨다. MCG- RTDETR은 UAV 영상에서의 극단적인 스케일 및 밀집 분포 상황을 고려하여, 다중 컨볼루션 (convolution)과 문맥 유도 (context-guided) 구조, 그리고 캐스케이드 그룹 어텐션 (cascade group attention)을 결합함으로써 미세 특징 보존, 문맥 기반 인코딩, 그리고 소형 객체 탐지 성능을 강화하는 동시에 실시간 적용 가능성을 유지한다. 이종 센싱 환경을 위해 제안된 교차 모달 (cross-modality) 특징 적응 탐지 프레임워크인 CMFADet은 RGB-적외선 및 SAR-광학 쌍 영상에 대해 이중 스트림 (dual-stream) 구조를 채택하고, 모달리티별 안정화, 경량 잔차 (residual) 상호작용, 그리고 적응적 과업 인지 정렬을 통합적으로 적용함으로써 표현 불일치와 분류 및 위치 추정 간의 불일치를 완화한다. 다양한 벤치마크 (benchmark) 데이터셋에 대한 광범위한 실험 결과는 제안된 방법론이 서로 다른 센싱 조건 전반에서 탐지의 강건성과 효율성을 일관되게 향상시킴을 입증한다. 구체적으로, SMEP-DETR은 SSDD 데이터셋에서 98.6% 𝑚𝐴𝑃 를 달성하였으며, MCG-RTDETR은 VisDrone2019에서 58.2% 𝐴𝑃50 을 기록하였다. 또한 CMFADet은 3.64M의 파라미터 수만으로 DroneVehicle 데이터셋에서 83.42% 𝑚𝐴𝑃50 을 달성하여, 정확도와 효율성 간의 우수한 균형을 보여준다. 이러한 결과는 센싱 유발 불일치를 확장 가능한 아키텍처 방법론을 통해 해결함으로써, 단일 모달리티 (single-modality) 및 교차 모달리티 (cross-modality) 원격 탐사 환경 전반에서 정확하고 실제 배치 가능한 객체 탐지가 가능함을 입증한다.
    번역하기

    학원 원격 탐사 (Remote Sensing, RS) 객체 탐지는 실제 센싱 환경에서 지속적인 어려움을 겪고 있으며, 그 주요 오류 원인은 네트워크 용량 자체보다는 센서 물리 특성과 취득 기하에 의해 지배된...

    학원 원격 탐사 (Remote Sensing, RS) 객체 탐지는 실제 센싱 환경에서 지속적인 어려움을 겪고 있으며, 그 주요 오류 원인은 네트워크 용량 자체보다는 센서 물리 특성과 취득 기하에 의해 지배된다. 합성 개구 레이다 (Synthetic Aperture Radar, SAR) 영상에서는 반점 노이즈 (Speckle)와 약한 객체 경계로 인해 클러터 기반 오경보 (false alarm) 와 위치 추정의 모호성이 발생한다. 무인항공기 (Unmanned Aerial Vehicle, UAV) 기반 광학 영상에서는 극소형 객체, 높은 공간 밀집도, 그리고 빈번한 가림 현상으로 인해 특징 표현이 심각하게 저하되며 탐지 실패가 증가한다. 또한 다중 센싱 모달리티 (Multi-modality sensing)를 공동으로 활용하는 경우, 서로 다른 영상 형성 메커니즘으로 인해 표현 불일치, 공간 매칭의 불완전함, 그리고 과업 수준의 불일치를 초래할 수 있으며, 단순한 초기 융합 전략은 불안정할 수 있고 심지어 성능 저하를 초래할 수 있다. 이러한 문제를 해결하기 위해 본 학위논문은 원격 탐사 객체 탐지를 위한 확장 가능하고 제약 기반의 고효율 아키텍처 방법론을 제안한다. 제안된 방법론은 탐지 성능을 모델 용량이나 융합 복잡도의 함수로 간주하는 기존 전통적인 방법과 달리, 센싱으로 유발된 불일치가 탐지 프로세스의 어느 단계에서 발생하는지를 명확하게 분석하고, 이를 세 단계를 통해 점진적으로 제거한다. 구체적으로, 모달리티별 표현 안정화, 제어된 경량 교차 모달리티 상호 작용 및 탐지 헤드에서의 과업 일관적 예측을 통해 불일치를 해결한다. 이러한 단계 인지적 (stage-aware) 설계는 단일 모달리티 최적화에서 교차 모달 통합으로 자연스럽게 확장 가능한 통합 설계 논리를 제공하며, 불필요한 계산 오버헤드 (overhead)를 초래하지 않는다. 제안된 방법론은 세 가지 대표적인 객체 탐지 프레임워크를 통해 구체화되고 검증된다. SMEP-DETR은 SAR 영상에 특화된 트랜스포머 기반 (transformer-based) 탐지 모델로서, 다중 에지 (edge) 강화와 병렬 팽창 문맥 집계를 통해 반점 노이즈 안정화와 명시적 경계 강화를 수행함으로써 해양 환경에서의 탐지 정확도를 향상시킨다. MCG- RTDETR은 UAV 영상에서의 극단적인 스케일 및 밀집 분포 상황을 고려하여, 다중 컨볼루션 (convolution)과 문맥 유도 (context-guided) 구조, 그리고 캐스케이드 그룹 어텐션 (cascade group attention)을 결합함으로써 미세 특징 보존, 문맥 기반 인코딩, 그리고 소형 객체 탐지 성능을 강화하는 동시에 실시간 적용 가능성을 유지한다. 이종 센싱 환경을 위해 제안된 교차 모달 (cross-modality) 특징 적응 탐지 프레임워크인 CMFADet은 RGB-적외선 및 SAR-광학 쌍 영상에 대해 이중 스트림 (dual-stream) 구조를 채택하고, 모달리티별 안정화, 경량 잔차 (residual) 상호작용, 그리고 적응적 과업 인지 정렬을 통합적으로 적용함으로써 표현 불일치와 분류 및 위치 추정 간의 불일치를 완화한다. 다양한 벤치마크 (benchmark) 데이터셋에 대한 광범위한 실험 결과는 제안된 방법론이 서로 다른 센싱 조건 전반에서 탐지의 강건성과 효율성을 일관되게 향상시킴을 입증한다. 구체적으로, SMEP-DETR은 SSDD 데이터셋에서 98.6% 𝑚𝐴𝑃 를 달성하였으며, MCG-RTDETR은 VisDrone2019에서 58.2% 𝐴𝑃50 을 기록하였다. 또한 CMFADet은 3.64M의 파라미터 수만으로 DroneVehicle 데이터셋에서 83.42% 𝑚𝐴𝑃50 을 달성하여, 정확도와 효율성 간의 우수한 균형을 보여준다. 이러한 결과는 센싱 유발 불일치를 확장 가능한 아키텍처 방법론을 통해 해결함으로써, 단일 모달리티 (single-modality) 및 교차 모달리티 (cross-modality) 원격 탐사 환경 전반에서 정확하고 실제 배치 가능한 객체 탐지가 가능함을 입증한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Remote sensing (RS) object detection faces persistent challenges under real- world sensing conditions, where dominant error sources are governed by sensing physics and acquisition geometry rather than network capacity alone. In synthetic aperture radar (SAR) imagery, coherent speckle and weak object boundaries induce clutter-driven false alarms and localization ambiguity. In unmanned aerial vehicle (UAV) optical imagery, extremely small objects, dense spatial distributions, and frequent occlusion lead to severe feature degradation and missed detections. When multiple sensing modalities are jointly exploited, heterogeneous imaging mechanisms further introduce representation mismatch, imperfect spatial correspondence, and task-level inconsistency, rendering early fusion strategies unstable or even detrimental. To address these challenges under strict efficiency constraints, this dissertation develops a scalable, constraint-driven architectural methodology for high-efficiency object detection in remote sensing. Rather than treating detection performance as a function of model capacity or fusion complexity, the proposed methodology explicitly investigates where sensing-induced discrepancy enters the detection pipeline and resolves it progressively through three stages: stabilizing modality- specific representations, enabling controlled and lightweight cross-modality interaction, and enforcing task-consistent prediction at the detection head. This stage-aware formulation establishes a unified design logic that generalizes from single-modality optimization to cross-modality integration without incurring unnecessary computational overhead. The proposed methodology is instantiated and validated through three representative detection frameworks. SMEP-DETR is a transformer-based detector for SAR imagery with multi-edge enhancement and parallel dilated context aggregation, which introduces speckle noise stabilization and edge enhancement to improve detection accuracy in maritime scenes. MCG-RTDETR, a multi- convolution and context-guided network with cascaded group attention for UAV imagery, is proposed to enhance fine-grained feature preservation, context-guided encoding, and small-object detection while maintaining real-time viability under extreme scale and density constraints. For heterogeneous sensing scenarios, a cross- modality feature adaptive detection framework, termed CMFADet, is proposed. CMFADet adopts a dual-stream architecture for paired RGB-infrared and SAR- optical imagery, where modality-specific stabilization, lightweight residual interaction, and adaptive task-aware alignment are jointly employed to mitigate representation incompatibility and the inconsistency of classification and localization tasks. Extensive experiments on widely used benchmarks demonstrate that the proposed methodology consistently improves both detection robustness and efficiency across diverse sensing conditions. Specifically, SMEP-DETR achieves 98.6% 𝑚𝐴𝑃 on SSDD, MCG-RTDETR attains 58.2% 𝐴𝑃50 on VisDrone2019, and CMFADet reaches 83.42% 𝑚𝐴𝑃50 on DroneVehicle with only 3.64M parameters, validating a favorable accuracy and efficiency trade-off. These results confirm that resolving sensing-induced discrepancy through a scalable architectural methodology enables accurate and deployable object detection across both single- modality and cross-modality remote sensing scenarios.
    번역하기

    Remote sensing (RS) object detection faces persistent challenges under real- world sensing conditions, where dominant error sources are governed by sensing physics and acquisition geometry rather than network capacity alone. In synthetic aperture rada...

    Remote sensing (RS) object detection faces persistent challenges under real- world sensing conditions, where dominant error sources are governed by sensing physics and acquisition geometry rather than network capacity alone. In synthetic aperture radar (SAR) imagery, coherent speckle and weak object boundaries induce clutter-driven false alarms and localization ambiguity. In unmanned aerial vehicle (UAV) optical imagery, extremely small objects, dense spatial distributions, and frequent occlusion lead to severe feature degradation and missed detections. When multiple sensing modalities are jointly exploited, heterogeneous imaging mechanisms further introduce representation mismatch, imperfect spatial correspondence, and task-level inconsistency, rendering early fusion strategies unstable or even detrimental. To address these challenges under strict efficiency constraints, this dissertation develops a scalable, constraint-driven architectural methodology for high-efficiency object detection in remote sensing. Rather than treating detection performance as a function of model capacity or fusion complexity, the proposed methodology explicitly investigates where sensing-induced discrepancy enters the detection pipeline and resolves it progressively through three stages: stabilizing modality- specific representations, enabling controlled and lightweight cross-modality interaction, and enforcing task-consistent prediction at the detection head. This stage-aware formulation establishes a unified design logic that generalizes from single-modality optimization to cross-modality integration without incurring unnecessary computational overhead. The proposed methodology is instantiated and validated through three representative detection frameworks. SMEP-DETR is a transformer-based detector for SAR imagery with multi-edge enhancement and parallel dilated context aggregation, which introduces speckle noise stabilization and edge enhancement to improve detection accuracy in maritime scenes. MCG-RTDETR, a multi- convolution and context-guided network with cascaded group attention for UAV imagery, is proposed to enhance fine-grained feature preservation, context-guided encoding, and small-object detection while maintaining real-time viability under extreme scale and density constraints. For heterogeneous sensing scenarios, a cross- modality feature adaptive detection framework, termed CMFADet, is proposed. CMFADet adopts a dual-stream architecture for paired RGB-infrared and SAR- optical imagery, where modality-specific stabilization, lightweight residual interaction, and adaptive task-aware alignment are jointly employed to mitigate representation incompatibility and the inconsistency of classification and localization tasks. Extensive experiments on widely used benchmarks demonstrate that the proposed methodology consistently improves both detection robustness and efficiency across diverse sensing conditions. Specifically, SMEP-DETR achieves 98.6% 𝑚𝐴𝑃 on SSDD, MCG-RTDETR attains 58.2% 𝐴𝑃50 on VisDrone2019, and CMFADet reaches 83.42% 𝑚𝐴𝑃50 on DroneVehicle with only 3.64M parameters, validating a favorable accuracy and efficiency trade-off. These results confirm that resolving sensing-induced discrepancy through a scalable architectural methodology enables accurate and deployable object detection across both single- modality and cross-modality remote sensing scenarios.

    더보기

    목차 (Table of Contents)

    • CHAPTER1 Introduction 1
    • 1.1 Background and Remote Sensing Platforms 1
    • 1.2 Characteristics of Remote Sensing Imagery 2
    • 1.3 Motivation and Research Objectives 3
    • 1.4 Main Contributions 5
    • CHAPTER1 Introduction 1
    • 1.1 Background and Remote Sensing Platforms 1
    • 1.2 Characteristics of Remote Sensing Imagery 2
    • 1.3 Motivation and Research Objectives 3
    • 1.4 Main Contributions 5
    • 1.5 Organization of the Dissertation 5
    • CHAPTER2 Preliminaries and Related Work 7
    • 2.1 Problem Definition and Task Formulation 7
    • 2.2 DL-Based Object Detection Paradigms 8
    • 2.2.1 Detection Pipelines and Modular Architecture 8
    • 2.2.2 CNN-Based Detection Paradigms 9
    • 2.2.3 Transformer-Based Detection Paradigms 10
    • 2.3 Object Detection in Remote Sensing Scenarios 11
    • 2.3.1 UAV-based Detection and Small Object Challenges 11
    • 2.3.2 SAR-based Ship Detection and Noise Interference 12
    • 2.3.3 Cross-modality Detection and Feature Fusion 12
    • 2.4 Summary and Research Gaps 14
    • CHAPTER3 Single-Modality Object Detection 15
    • 3.1 Scope and Challenges of Single-Modality Detection 15
    • 3.2 SAR Ship Detection 16
    • 3.2.1 SAR Imaging Constraints 16
    • 3.2.2 Detection Architecture under SAR Constraints 17
    • 3.2.3 Feature Adaptation for Speckle and Boundary Ambiguity 18
    • 3.2.4 Uncertainty-Guided Query Selection 22
    • 3.3 UAV Small Object Detection 24
    • 3.3.1 Scale and Density Constraints 24
    • 3.3.2 Detection Architecture under UAV Constraints 25
    • 3.3.3 Feature Representation Enhancement 26
    • 3.3.4 Context-Guided Feature Encoding 28
    • 3.3.5 Prediction Design for Small Objects 30
    • 3.3.6 IoU-Aware Query Selection Strategy 31
    • 3.4 Comparative Analysis and Modality Limitations 33
    • CHAPTER4 Cross-Modality Object Detection 35
    • 4.1 Scope and Challenges of Cross-Modality Detection 35
    • 4.2 Cross-Modality Detection Framework 38
    • 4.2.1 Problem Decomposition under Cross-Modality Perception 39
    • 4.2.2 Overall Architectural Design Rationale 41
    • 4.2.3 Dual-Stream Backbone Network 42
    • 4.3 Spatial-Frequency Feature Enhancement Module 45
    • 4.4 Infrared Adaptive Feature Aggregation Block 47
    • 4.5 Channel Interaction Fusion 50
    • 4.6 Adaptive Task-Aware Alignment Head 51
    • 4.7 Multi-Task Objective Function 53
    • CHAPTER5 Experimental Evaluation and Analysis 55
    • 5.1 Experimental Protocol and Implementation Details 55
    • 5.2 Datasets and Benchmark Description 58
    • 5.2.1 SAR Ship Detection Datasets 58
    • 5.2.2 UAV Object Detection Datasets 60
    • 5.2.3 Cross-Modality Detection Datasets 61
    • 5.3 Evaluation Metrics 63
    • 5.4 Evaluation on Single-Modality Detection 66
    • 5.4.1 Quantitative Comparison with State-of-the-Art 66
    • 5.4.2 Qualitative Analysis and Visualization 72
    • 5.4.3 Ablation and Component-wise Analysis 82
    • 5.5 Evaluation on Cross-Modality Detection 87
    • 5.5.1 Quantitative Comparison with State-of-the-Art 87
    • 5.5.2 Qualitative Analysis and Visualization 93
    • 5.5.3 Ablation and Component-wise Analysis 105
    • CHAPTER6 Conclusion 109
    • 6.1 Overall Summary and Main Contributions· 109
    • 6.2 Discussion 112
    • 6.3 Future Work 114
    • REFERENCES 115
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼