최근 원격 탐사 기술의 발전으로 고해상도 항공 이미지의 접근성이 높아짐에 따라, 이를 분석하기 위한 딥러닝 기반의 의미론적 분할(Semantic Segmentation) 연구가 활발히 진행되고 있다. 초기에...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17407147
순천 : 국립순천대학교 대학원, 2026
학위논문(석사) -- 국립순천대학교 대학원 , 멀티미디어공학과 , 2026. 2
2026
한국어
전라남도
; 26 cm
지도교수: 심춘보
I804:46008-000000010963
0
상세조회0
다운로드최근 원격 탐사 기술의 발전으로 고해상도 항공 이미지의 접근성이 높아짐에 따라, 이를 분석하기 위한 딥러닝 기반의 의미론적 분할(Semantic Segmentation) 연구가 활발히 진행되고 있다. 초기에...
최근 원격 탐사 기술의 발전으로 고해상도 항공 이미지의 접근성이 높아짐에 따라, 이를 분석하기 위한 딥러닝 기반의 의미론적 분할(Semantic Segmentation) 연구가 활발히 진행되고 있다. 초기에는 CNN 기반 모델이 주로 활용되었으나, 고정된 수용 영역으로 인해 전역적 문맥(Global Context)을 포착하는 데 한계를 보였다. 이를 극복하기 위해 등장한 Transformer 기반 모델은 성능은 우수하나, 입력 시퀀스 길이에 따른 2차 계산 복잡도 문제로 막대한 연산 비용이 소모되는 단점이 있다. 이러한 문제를 해결하기 위해 선형적 계산 복잡도를 가진 Mamba 아키텍처가 주목받고 있다. 이를 비전 분야에 적용한 대표적인 시각적 상태 공간 모델인 VMamba는 전역적 스캔 방식에 의존하기 때문에 지역적 세부사항을 포착하는 데 한계가 있다. 특히 광활한 배경과 소형 객체가 혼재된 원격 탐사 이미지에서 소형 객체의 특징이 희석되거나 복잡한 경계선이 뭉개지는 문제가 발생한다. 본 논문에서는 VMamba의 구조적 한계를 극복하고 원격 탐사 이미지의 분할 정밀도를 향상시키기 위해 Shifted Window를 결합한 새로운 시각적 상태 공간 모델을 제안한다. W-SS2D(Window-based 2D-Selective Scan)와 SW-SS2D(Shifted Window-based 2D-Selective Scan)를 교차로 수행하는 이중 구조를 가진다. W-SS2D는 이미지를 윈도우 단위로 분할하여 스캔 범위를 지역적으로 제한함으로써 소형 객체와 미세한 디테일 포착 능력을 강화한다. 이어지는 SW-SS2D는 윈도우를 순환 이동(Cyclic Shift)시켜 단절되었던 윈도우 간의 정보 교류를 활성화하고 전역적 문맥 연결성을 확보한다. 특히, 본 연구는 윈도우 순환 이동 시 별도의 마스킹(Masking) 연산 없이도 S6 연산을 효율적으로 수행하도록 설계되었다. 디코더는 UPerNet 구조를 채택하여 다중 스케일 정보를 효과적으로 융합한다. 제안하는 모델의 성능을 검증하기 위해 다양한 지리적 환경을 포함한 LoveDA 데이터셋을 사용하여 실험을 진행했다. 실험 결과, 제안하는 모델은 mIoU 54.22%, mF1 69.51%를 달성하며 CNN 기반 모델뿐만 아니라 최신 Mamba 기반 모델보다 우수한 성능을 기록했다. 특히 건물(Building), 수역(Water), 황무지(Barren), 숲(Forest) 클래스에서 타 모델 대비 가장 높은 IoU를 기록했다. 정성적 평가 결과에서도 객체 간의 경계 흐림(Boundary Blurring) 현상이 현저히 감소하고 소형 객체를 정밀하게 분할함을 확인했다.
다국어 초록 (Multilingual Abstract)
With recent advancements in remote sensing technology increasing the accessibility of high-resolution aerial imagery, research on deep learning-based semantic segmentation for its analysis is being actively conducted. Initially, CNN-based models were ...
With recent advancements in remote sensing technology increasing the accessibility of high-resolution aerial imagery, research on deep learning-based semantic segmentation for its analysis is being actively conducted. Initially, CNN-based models were primarily used; however, they showed limitations in capturing global context due to their fixed receptive fields. Transformer-based models, introduced to overcome this, demonstrate superior performance but suffer from the disadvantage of incurring massive computational costs due to quadratic computational complexity relative to the input sequence length. To address these issues, the Mamba architecture, characterized by linear computational complexity, has attracted significant attention. VMamba, a representative visual state space model adapting Mamba for computer vision, relies on a global scanning mechanism, which limits its ability to effectively capture local details. This limitation is particularly pronounced in remote sensing imagery where vast backgrounds and small-scale objects coexist, often resulting in diluted features of small objects and blurred complex boundaries. To overcome the structural limitations of VMamba and improve segmentation precision in remote sensing imagery, this paper proposes a novel visual state space model incorporating Shifted Window. The proposed model features a structure that alternately performs W-SS2D (Window-based 2D-Selective Scan) and SW-SS2D (Shifted Window-based 2D-Selective Scan). W-SS2D partitions the image into windows to locally limit the scanning range, thereby enhancing the capability to capture small objects and fine details. The subsequent SW-SS2D employs cyclic shifting of windows to activate information exchange between disconnected windows and ensure global context connectivity. Notably, the proposed method is designed to efficiently perform S6 operations without requiring additional masking operations during window shifting. The decoder adopts the UPerNet architecture to effectively fuse multi-scale information. Experiments were conducted using the LoveDA dataset, which encompasses diverse geographical environments, to validate the performance of the proposed model. Experimental results demonstrate that the proposed model achieved an mIoU of 54.22% and an mF1 of 69.51%, outperforming not only CNN-based models but also state-of-the-art Mamba-based models. In particular, it recorded the highest IoU scores in the Building, Water, Barren, and Forest classes. Qualitative evaluations also confirmed that boundary blurring between objects was significantly reduced, and small objects were segmented with high precision.
목차 (Table of Contents)