RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    검색결과 좁혀 보기

    선택해제
    • 좁혀본 항목 보기순서

      • 원문유무
      • 음성지원유무
      • 학위유형
      • 주제분류
      • 수여기관
        펼치기
      • 발행연도
      • 작성언어
      • 지도교수
        펼치기

    오늘 본 자료

    • 오늘 본 자료가 없습니다.
    더보기
    • Swin Transformer와 Cascade R-CNN 혼합 모형을 활용한 소나무 병해충 탐지 시스템에 관한 연구

      가오가오 호남대학교 대학원 2022 국내석사

      RANK : 232319

      소나무과에 속하는 소나무는 한국의 산림에 가장 넓게 분포되어 있으며 개체 수 역시 다른 산림 수보다 많으며, 활용 용도에 따라 건축, 전신주, 교량, 농기구, 기구, 가구 그리고 제작제지업과 귀한 약재 원료로 사용하고 있으며, 솔가지와 솔뿌리는 먹, 잉크, 검은색 도색 재료 등 다양한 재료에 쓰인다. 이러한 소나무는 다른 나무들에 비해 병해충에 매우 취약한 면을 보이고 있으며, 병해충 종류로는 솔잎 흑파리, 솔껍질깍지벌레, 참나무시들음병 등이 있다. 현재 소나무 병해충을 확인하기 위해 산림관리자가 직접 수작업으로 채집망을 이용해 채집하여 눈으로 확인하는 과정을 진행하고 있으므로 빠른 방제작업이 늦어지는 원인이 된다. 따라서 본 논문은 딥러닝을 이용하여 소나무 병해충을 빠른 시간에 확인할 수 있는 소나무 병해충 탐지 시스템을 구현하고자 한다. 또한, 인공지능의 좋은 성능 모형을 선택하여 시스템에 적용하기 위해 You Only Look Once (YOLOv5s)_Focus+C3, Cascade Region-based Convolutional Neural Networks (Cascade R-CNN)_Residual Network 50(ResNet50), Faster Region-based Convolutional Neural Networks (Faster R-CNN)_ResNet50 그리고 본 논문에서 제안된 Swin Transformer와 Cascade R-CNN의 혼합 모형 등의 4가지 성능을 비교 분석하고자 한다. 결과는 YOLOv5s_Focus+C3은 소나무 병해충 탐지에서 Average Precision (AP) Intersection over Union (IoU=0.5)는 56.3%, Recall 값은 66.8%로 나타났으며, Cascade R-CNN_ResNet50은 소나무 병해충 탐지에서 AP(IoU=0.5) 93.2%, AP(IoU=0.75) 89.5%, Recall 92.9%로 나타났다. Cascade R-CNN_ResNet50은 Faster R-CNN보다 AP(IoU=0.5)는 0.7% 높고, AP(IoU=0.75)는 1.6% Recall은 1.8%로 높은 값을 가진다. 반면 Swin Transformer와 Cascade R-CNN의 혼합 모형은 Cascade R-CNN_ResNet50 과 비교해 AP(IoU=0.5)가 1.4%만큼 높으며, AP(IoU=0.75)는 0.1%만큼 높다. 또한, Recall은 0.6%만큼 높게 나타났다. 따라서 소나무 병해충 탐지하는 4가지 모형의 성능을 비교한 결과 YOLOv5s_Focus+C3, Faster R-CNN_ResNet50, Cascade R-CNN_ResNet50, 그리고 본 논문에서 제안하는 Cascade R-CNN_Swin Transformer 혼합 모형과 비교하면 낮은 성능을 보였다. 따라서 본 논문에서 제안한 Swin Transformer와 Cascade R-CNN의 혼합 모형이 비교한 4가지 모형 중 우수함이 본 연구를 통해 확인되었다. Pine trees, which belong to the Pinaceae, are the most broadly distributed in the forests of Korea, and have a higher population than other trees. Their diverse applications include construction materials, telephone poles, bridges, tools, furniture, papermaking, and medicinal ingredients. Pine branches and roots are also used to make inks and black coloring materials. However, pine trees are more vulnerable to pests, with common types being the pine needle gall midge, black pine bast scale, and Korean oak wilt. Currently, forest managers manually collect samples and inspect them for larvae using the naked eye. This is a factor contributing to the slow rate of pest control. Against this backdrop, this study seeks to develop a deep learning-based pine larva detection system to allow more rapid and effective pest control. To select a model with the best performance, a comparative analysis was carried out on four models: You Only Look Once (YOLOv5s)_Focus+C3, Cascade Region-based Convolutional Neural Networks (Cascade R-CNN)_Residual Network 50(ResNet50), Faster Region-based Convolutional Neural Networks (Faster R-CNN)_ResNet50, and the proposed Swin Transformer and Cascade R-CNN hybrid model. The analysis revealed that YOLOv5s_Focus+C3 had an Average Precision (AP) Intersection over Union (IoU=0.5) of 56.3% and a Recall of 66.8% in pine larva detection. Cascade R-CNN_ResNet50 had an AP (IoU=0.5) of 93.2%, an AP (IoU=0.75) of 89.5%, and a Recall of 92.9%. Compared to Faster R-CNN, Cascade R-CNN_ResNet50 had higher AP(IoU=0.5) by 0.7%, AP(IoU=0.75) by 1.6%, and Recall by 1.8%. Compared to Cascade R-CNN_ResNet50, theproposed Swin Transformer and Cascade R-CNN hybrid model had higher AP(IoU=0.5) by 1.4% and higher AP(IoU=0.75) by 0.1%. In addition, its recall was higher by 0.6%. Based on the above comparison of the four models of pine larva detection, YOLOv5s_Focus+C3, Faster R-CNN_ResNet50, and Cascade R-CNN_ResNet50 yielded poorer performance than the proposed Cascade R-CNN_Swin Transformer hybrid model. That is, the proposed hybrid model of Swin Transformer and Cascade R-CNN was found to be the most outstanding among the four models.

    • Residual Dense Network와 SEBlock을 활용한 Swin Transformer 기반 초해상도 이미지 복원 성능 향상

      김찬민 서강대학교 일반대학원 2025 국내석사

      RANK : 232317

      본 연구는 Swin Transformer를 기반으로 한 새로운 초해상도 이미지 복원 모델을 제안하여, 복잡한 텍스처와 세부 정보를 효과적으로 복원하는 데 기여한다. 최근 딥러닝 기반의 이미지 초해상도(Super-Resolution, SR) 기술은 의료 영상, 위성 사진, 감시 시스템 등 다양한 분야에서 저해상도 이미지를 고해상도로 복원하여 활용성을 높이고 있다. 기존의 SwinIR 모델은 Swin Transformer 블록을 활용하여 전역적인 특징 학습에 우수한 성능을 보였으나, 지역적인 세부 정보 추출과 채널 간 상호 의존성 학습에는 한계가 있었다. 본 연구에서는 SwinIR 모델의 성능 향상을 위해 Residual Dense Network(RDN)와 Squeeze-and-Excitation Block(SEBlock)을 효과적으로 통합한 새로운 신경망을 제안한다. 첫째, 얕은 특징 추출 단계에서 RDN을 적용하여 지역적인 특징 학습을 강화한 RSwinIR 모델을 제안한다. 둘째, Residual Swin Transformer Block(RSTB) 내부에 SEBlock을 삽입하여 채널 어텐션 메커니즘을 도입한 SESwinIR 모델을 설계하였다. 마지막으로, RDN과 SEBlock을 모두 적용하여 모델의 성능을 극대화한 RSE-SwinIR 모델을 제안하였다. 실험은 Set5, Set14, BSD100 데이터셋에서 수행되었으며, 제안된 모델들은 기존 SwinIR 대비 PSNR 및 SSIM 지표에서 우수한 성능을 보였다. 특히 RSE-SwinIR 모델은 이미지 구조 복원 성능의 향상을 보여주어, RDN과 SEBlock의 결합이 이미지 초해상도 작업에서 효과적임을 입증하였다. This study proposes a novel super-resolution image restoration model based on the Swin Transformer, contributing to the effective reconstruction of complex textures and fine details. Recent advances in deep learning-based image super-resolution (SR) technology have significantly enhanced the restoration of high-resolution images from low- resolution inputs in various fields such as medical imaging, satellite photography, surveillance systems, and video streaming services. The existing Swin Transformer for Image Restoration (SwinIR) model excels in global feature learning using Swin Transformer blocks but has limitations in extracting local fine details and learning inter-channel dependencies. In this study, we propose a new neural network that effectively integrates the Residual Dense Network (RDN) and the Squeeze-and-Excitation Block (SEBlock) to enhance the performance of the SwinIR model. First, we introduce the RSwinIR model by applying RDN to the shallow feature extraction stage to strengthen local feature learning. Second, we design the SESwinIR model by inserting SEBlock into the Residual Swin Transformer Block (RSTB) to introduce a channel attention mechanism. Finally, we propose the RSE- SwinIR model that maximizes performance by applying both RDN and SEBlock. Experiments were conducted on Set5, Set14, and BSD100 datasets. The proposed models demonstrated superior performance in PSNR and SSIM metrics compared to the existing SwinIR model. In particular, the RSE-SwinIR model showed improvements in image structure restoration performance, demonstrating that the combination of RDN and SEBlock is effective in image super-resolution tasks.

    • Automatic Modulation Recognition of Swin Transformer Using Multimodal Information

      Shin, Da Min 부산대학교 대학원 2026 국내석사

      RANK : 232316

      현대 전자전 환경에서는 통신 기술의 고도화로 인해 전장에서 관측되는 신호양과 형태가 급격히 증가하고 있으며, 이로 인해 수신 신호를 정확히 분석하는 일이 점점 더 어려워지고 있다. 또한 민수 분야의 무선 통신 분야에서는 수신 신호의 변조 기법을 분석하여 스펙트럼 모니터링과 같은 분야에 사용하고자 한다. 이러한 배경에서 다양한 통신 신호를 높은 정확도로 식별하기 위한 딥러닝 기반 자동 변조 인식 기법이 활발히 연구되고 있다. 본 논문에서는 1차원 Swin Transformer 구조를 토대로 한 다중 모달리티 자동 변조 인식 신경망을 제안한다. 1D Swin Transformer는 계층적으로 구성된 특징 맵과 이동 윈도우 기반 self-attention을 활용하여, 입력 신호로부터 서로 다른 해상도의 유용한 특징을 효율적으로 추출한다. 제안하는 네트워크는 IQ 시계열 신호와 주파수 영역 스펙트럼을 동시에 입력으로 사용하는 다중 모달 구조로 설계되었으며, 각 모달리티에서 개별적으로 특징을 추출한 뒤 이를 1D Swin Transformer에 결합하여 통합 표현을 형성하고, 이를 바탕으로 변조 방식을 분류한다. 실험 결과, 제안한 모델은 평균 61.45%의 분류 정확도를 달성하여 비교 대상으로 설정한 기존 신경망들보다 우수한 성능을 보였다. 특히 QAM 계열 신호에 대해서는 평균 58.12%의 인식 정확도를 기록함으로써, 파형 구조가 유사한 변조 방식들 사이에서도 효과적인 구분 능력을 지니고 있음을 확인하였다. In modern electronic warfare environments, advances in communication technology have led to a rapid increase in both the volume and diversity of signals observed on the battlefield, making accurate analysis of received signals increasingly challenging. In addition, in civilian wireless communication systems, analyzing the modulation schemes of received signals is required for applications such as spectrum monitoring. Against this background, deep learning–based automatic modulation recognition (AMR) techniques that can accurately identify various communication signals have been actively studied. In this thesis, we propose a multi-modal automatic modulation recognition neural network based on a one-dimensional (1D) Swin Transformer architecture. The 1D Swin Transformer exploits hierarchically structured feature maps and a shifted-window self-attention mechanism to efficiently extract informative features at multiple resolutions from the input signal. The proposed network adopts a multi-modal structure that takes both IQ time-series signals and frequency-domain spectra as inputs. Modality-specific feature extractors are first applied to each input, and the resulting features are then fused and fed into the 1D Swin Transformer to obtain an integrated representation, which is finally used to classify the modulation scheme. Experimental results show that the proposed model achieves an average classification accuracy of 61.45%, outperforming the baseline neural networks considered for comparison. In particular, it achieves an average accuracy of 58.12% for QAM-family signals, demonstrating strong discriminative capability even among modulation types with similar waveform structures.

    • Fusion Network of CNN and Swin Transformer for Small Target Detection in Infrared Images

      최람미 서울대학교 대학원 2024 국내석사

      RANK : 232315

      Fusion Network of CNN and Swin Transformer for Small Target Detection in Infrared Images Lammi Choi Department of Aerospace Engineering The Graduate School Seoul National University In the realm of infrared (IR) small target detection, accurately identifying indistinct and low-contrast targets poses a formidable challenge owing to the intricate and cluttered characteristics of IR images. This study introduces a network architecture termed the CNN-Swin Transformer integrated network (CSI-Net) to tackle this challenge. The proposed network incorporates a hybrid encoder structure that fuses the encoder-decoder design of UNet with the parallel execution of Swin Transformer alongside CNN. This amalgamation enables the network to adeptly capture both local features and long-range dependencies, augmenting its ability to precisely discern small targets. By leveraging the hierarchical features of the Swin Transformer, the CSI-Net adeptly captures contextual information and intricate patterns crucial for small target detection. Moreover, the CSI-Net employs a full- scale skip connection across encoder-decoder and decoder-decoder, seamlessly integrating multiscale CNN and Swin Transformer features, thereby enhancing gradient propagation. Experimental findings unequivocally showcase the superiority of the proposed CSI-Net over conventional CNN and Transformer approaches, as evidenced by improvements in mean Intersection over Union (mIoU), probability of detection, and false alarm rate. This comprehensive evaluation attests to the efficacy of the Swin Transformer-enhanced UNet architecture in effectively addressing the challenges associated with IR small target detection. Keywords: Target detection, Infrared (IR) image, small target, Swin Transformer, UNet, hybrid encoder Student number: 2022-26314 적외선 소형 목표물 탐지(Infrared small target detection, IRSTD) 분야에서는 흐릿하고 낮은 대비의 목표물을 정확하게 식별하는 것에서 적외선 이미지의 복잡하고 혼잡한 특징으로 인해 어려움을 겪고 있다. 이러한 도전에 대응하기 위해 본 연구에서는 CNN-Swin Transformer 통합 네트워크(CSINet)라는 네트워크 아키텍처를 제안하였다. CSI-Net은 기존의 UNet의 인코더-디코더 디자인과 Swin Transformer의 병렬 실행을 결합한 하이브리드 인코더 구조를 도입하여, 소규모 목표물을 더욱 정확하게 식별할 수 있는 능력을 향상시켰다. 이러한 조합은 네트워크가 지역적인 특징과 장거리 종속성을 동시에 포착할 수 있도록 하여 전체적인 목표물 감지 성능을 향상시킨다. CSI-Net은 Swin Transformer의 계층적 특징을 활용하여 목표물 탐지에 필수적인 맥락 정보와 복잡한 패턴을 효과적으로 이해할 수 있다. 또한, 본네트워크는 인코더-디코더 간 및 디코더-디코더 간의 full-scale skip connection을 적용하여 다중 스케일 CNN과 Swin Transformer 특징을 효과적으로 통합하고, gradient 전파를 향상시킨다. 실험 결과는 CSI-Net이 기존의 CNN 및 Transformer 방법에 비해 mIoU, 감지 확률 (probability of detection) 및 거짓 경보율 (false alarm rate) 측면에서 우수한 성능을 보여준다. 이러한 종합적인 평가 결과로 보아, Swin Transformer가 결합된 UNet 아키텍처를 통해 적외선 소형 목표물 탐지에서의 어려움을 개선할 수 있음을 확인할 수 있었다.

    • On the head redundancy in Swin transformer for image classification

      오준호 Graduate School, Yonsei University 2022 국내석사

      RANK : 232287

      Transformer 모델은 원래 자연어 처리를 위해 고안된 모델이지만, computer vision 등의 타 분야에서도 Transformer를 활용하려는 연구가 매우 활발하게 진행되고 있다. 자연어 처리 분야에서 이용되는 Transformer 기반 모델의 Multi-Head Self-Attention (MHSA) 모듈을 구성하는 각 헤드간의 중복성에 관한 연구는 존재하지만, computer vision 분야에서 이용되는 Transformer 기반 모델에 대해 이를 연구한 바는 없다. 본 논문에서는 Swin Transformer 모델의 MHSA 모듈에서 각 헤드를 제거했을 때 ImageNet 데이터셋의 이미지 분류 작업에서 모델 성능의 변화를 측정하였고, 이를 통해 몇몇 헤드는 다른 헤드와 중복성이 존재하여, 그 헤드를 제거하더라도 정확도의 하락이 적거나 오히려 상승함을 밝혔다. 또한 stage 3을 구성하는 헤드 중 절반을 제거하여 비교적 적은 정확도 감소를 대가로 모델을 경량화할 수 있음을 보였다. 마지막으로, 헤드의 중복성에 영향을 미칠 것으로 예상되는 3개의 잠재요소로 출력 행렬의 노름의 평균값, 입력에 따른 출력 행렬 간의 불변성, 각 헤드 간 attention map의 코사인 유사도를 제시하였고, 그 중 전자 2개의 요소와 모델 성능 간에 상관관계가 존재함을 발견하였다. Although the Transformer model was originally designed for natural language processing (NLP), many studies are being actively conducted in utilizing the Transformer in other fields such as computer vision. There is a study on the redundancy between attention heads in multi-head self-attention (MHSA) modules of Transformer-based models in NLP, but there is no study on Transformer-based models in computer vision. In this thesis, we measure the change of model performance on the ImageNet image classification task when each head in MHSA layers in Swin Transformer is removed. Several heads are redundant, so the model accuracy slightly decreases or rather increases. In addition, it is shown that the model can be compressed by removing half of the heads in stage 3, in exchange for insignificant accuracy loss. Finally, we offer three factors that are expected to affect the head redundancy; the mean of norms of output matrices for various inputs, the invariance of output matrices for various inputs, and the cosine similarity between attention maps of each head. It is proved that the former two factors are correlated with head redundancy.

    • UDLV3+를 활용한 인프라 구조물 균열 탐지 설계 및 성능 평가

      이종현 호남대학교 대학원 2025 국내박사

      RANK : 232283

      Design and Performance Evaluation of Infrastructure Crack Detection Using UDLV3+ Lee Jong-Hyun Department of Computer Engineering Honam University Directed by prof. Lee Sang-Hyun This study aims to design and evaluate the performance of deep learning-based segmentation models for more precise and efficient crack detection in concrete-based infrastructure structures. In infrastructure maintenance, cracks are among the most critical defects, and accurately detecting early-stage cracks plays a key role in determining the overall safety and lifespan of the structure. However, conventional visual inspection methods and crack gauge-based measurements have clear limitations in terms of objectivity and quantifiability, and are prone to subjective inconsistencies depending on the inspector’s expertise. To address these issues, this research proposes an automated crack detection method using artificial intelligence technology, specifically image-based deep learning segmentation models. To achieve this, four models were selected for comparative evaluation: the widely adopted U-Net, DeepLabV3+, Swin Transformer, and the proposed hybrid model combining U-Net and DeepLabV3+ (UDLV3+). The models were compared based on their architectural features, training stability, inference speed, and quantitative performance indicators. The proposed UDLV3+ model was designed to combine the lightweight structure and fast convergence of U-Net with the multi-scale feature extraction and boundary refinement capabilities of DeepLabV3+, aiming to maximize crack detection accuracy while maintaining computational efficiency suitable for real-time applications. Experiments were conducted using a dataset of over 40,000 concrete surface images classified into Positive (with cracks) and Negative (without cracks) categories. Preprocessing and augmentation techniques such as mask generation, normalization, resizing, and horizontal flipping were applied to improve generalization performance during training. The models were trained for 50 epochs, and performance was evaluated using metrics including IoU, Dice coefficient, inference speed (FPS), computational complexity (GFLOPs), and the number of parameters (Million Parameters). As a result, the proposed UDLV3+ model achieved the highest detection performance with an IoU of 0.523 and a Dice coefficient of 0.554. Furthermore, it maintained a computational cost of only 1.43 GFLOPs and 26.18M parameters while achieving an inference speed of 519.70 FPS, demonstrating a successful balance between accuracy and efficiency. The U-Net model, with a Dice coefficient of 0.472 and an inference speed of 1978.91 FPS, showed extremely fast processing capabilities and was found to be suitable for lightweight or edge device-based systems. In contrast, DeepLabV3+ effectively captured cracks of varying shapes and scales, but its high computational cost (10.26 GFLOPs) and large parameter count (39.63M) make it less suitable for resource-constrained environments. Despite being a state-of-the-art Vision Transformer-based architecture, the Swin Transformer showed poor accuracy with an IoU of 0.125 and Dice coefficient of 0.216, indicating difficulty in distinguishing cracks from background noise and blurry boundaries. Additionally, analysis of training loss curves showed that both the U-Net and UDLV3+ models exhibited rapid loss reduction in early epochs followed by a stable convergence phase. The hybrid model, in particular, demonstrated both fast convergence and low final loss values, confirming its robustness in training and strong generalization potential. In conclusion, the proposed UDLV3+ model outperformed individual models in terms of accuracy, computational efficiency, and inference speed, demonstrating a level of performance suitable for practical application in real-world crack diagnostic systems. Furthermore, this study presents a viable AI-based solution for various application environments such as automated infrastructure maintenance, unmanned inspection systems, and Edge-AI-based real-time diagnostics. These findings suggest a foundational contribution to the advancement of smart city infrastructure management and AI-integrated structural health monitoring systems. UDLV3+를 활용한 인프라 구조물 균열 탐지 설계 및 성능 평가 제 출 자 : 이 종 현 지도교수 : 이 상 현 본 논문은 콘크리트 기반 인프라 구조물에서 발생하는 균열을 보다 정밀하고 효율적으로 탐지하기 위한 딥러닝 기반 세그멘테이션 모델의 설계 및 성능 평가를 목적으로 한다. 구조물의 유지보수에서 균열은 주요 결함 중 하나이며, 균열의 초기 발생을 정확하게 감지하는 것은 전체 구조물의 안정성과 수명을 판단하는 데 핵심적인 역할을 한다. 그러나 기존의 시각 점검 방식이나 균열 게이지를 활용한 방법은 정량성과 신뢰성에서 뚜렷한 한계를 가지며, 전문가에 따른 주관적 편차 또한 발생할 수 있다. 이러한 문제를 해결하기 위해 본 연구는 인공지능 기술, 특히 영상 기반 딥러닝 세그멘테이션 모델을 활용하여 자동화된 균열 탐지 방법을 제안하고자 한다. 이를 위해 대표적인 세그멘테이션 모델인 U-Net, DeepLabV3+, Swin Transformer, 그리고 제안된 하이브리드 구조인 U-Net + DeepLabV3+ (UDLV3+)를 실험에 포함하여 각 모델의 구조적 특성, 학습 안정성, 추론 속도 및 정량 평가 지표를 기반으로 성능을 비교하였다. 특히 제안된 UDLV3+ 모델은 U-Net 모델의 경량성과 빠른 수렴 특성, DeepLabV3+ 모델의 다중 해상도 특징 추출 능력 및 경계 인식 강점을 결합하여, 균열 탐지 정확도를 극대화하면서도 실시간 응용에 적합한 연산 효율을 구현하는 데 목적을 두었다. 실험은 Positive(균열 포함)와 Negative(정상 이미지)로 구분된 40,000장 이상의 데이터셋을 기반으로 수행되었으며, 이미지 전처리 과정에서는 마스크 생성, 정규화, 크기 조정, 수평 반전 등의 데이터 증강 기법을 적용하여 학습의 일반화 성능을 확보하였다. 학습은 총 50 Epoch 동안 수행되었고, 성능 평가지표는 IoU, Dice 계수, 추론 속도(FPS), 연산량(GFLOPs), 파라미터 수(Million parameters)를 채택하였다. 구현 결과, 제안된 UDLV3+ 모델은 IoU 0.523, Dice 계수 0.554를 기록하여 탐지 정확도 면에서 가장 우수한 성능을 보였다. 해당 모델은 또한 1.43 GFLOPs의 연산량과 26.18M의 파라미터 수를 유지하면서도 519.70 FPS의 속도를 기록해, 정확도와 연산 효율 간의 균형을 성공적으로 달성한 것으로 나타났다. U-Net 단일 모델은 Dice 계수 0.513 및 FPS 1,978.91로 매우 빠른 추론 속도를 보이며, 경량 환경이나 엣지 디바이스 기반 시스템에 적합한 모델로 평가되었다. DeepLabV3+ 모델은 다양한 크기와 형태의 균열을 효과적으로 포착할 수 있으나, 10.26 GFLOPs의 높은 연산량과 39.63M의 파라미터 수로 인해 경량 환경에는 부적합한 것으로 나타났다. Swin Transformer 모델은 Vision Transformer 기반의 최신 구조임에도 불구하고 IoU 0.125, Dice 계수 0.216으로 정확도가 낮았으며, 균열과 배경 간 경계 식별에 어려움을 겪는 것으로 분석되었다. 또한 학습 중 손실(Loss) 변화 곡선 분석에서도 U-Net 및 UDLV3+ 모델은 초반 Epoch에서 손실이 급격히 감소한 후, 후반부에 걸쳐 안정적인 수렴 양상을 보였다. 특히 UDLV3+ 모델은 빠른 손실 감소와 낮은 최종 손실 값을 함께 보여, 모델 학습 안정성과 일반화 가능성 면에서도 우수함을 입증하였다. 따라서, 본 연구에서 제안한 UDLV3+ 모델은 기존 단일 모델 대비 정확도, 효율성, 추론 속도 등 다양한 측면에서 뛰어난 성능을 입증하였으며, 이는 실제 인프라 구조물의 균열 진단 시스템에 실용적으로 적용이 가능한 수준의 품질을 갖춘 것으로 판단된다. 또한 본 연구는 구조물의 유지보수 자동화, 무인 점검 시스템, Edge-AI 기반 실시간 진단 등 다양한 응용 환경에 적합한 AI 기반 균열 탐지 솔루션을 제시함으로써, 향후 스마트시티 및 인프라 관리 기술 고도화에 기여할 수 있는 기반이 될 수 있다.

    • Hybrid CNN?Swin Transformer Denoising Scheme for Channel Estimation in Optical IRS-Assisted Indoor MIMO VLC System

      Rafat Bin Mofidul 국민대학교 일반대학원 2025 국내석사

      RANK : 232269

      This research investigates advanced methods for estimating channel states in indoor MIMO Visible Light Communication (VLC) networks augmented with Optical Intelligent Reflective Surfaces (OIRS). Although VLC offers substantial bandwidth and immunity to electromagnetic interference, obtaining precise channel information is difficult. These difficulties stem from multipath signal reflections, non-line-of-sight (NLOS) paths, and intricate noise patterns. To overcome these obstacles, we introduce HCSTNet, a novel hybrid deep learning architecture. This framework combines a Convolutional Neural Network (CNN) dedicated to estimating noise with a Swin Transformer module designed for denoising. By merging these components, the system effectively captures both local noise details and long-range spatial correlations within the channel data. We established a comprehensive simulation environment modeling an indoor MIMO VLC setup that accounts for line-of-sight (LOS), NLOS, and OIRS-reflected paths. The training data includes realistic impairments such as thermal and shot noise. Experimental outcomes indicate that HCSTNet outperforms existing traditional and deep learning approaches. Specifically, the model reduces the Normalized Mean Square Error (NMSE) to the range of 10^-3 to 10^-4 while significantly boosting Peak Signal-to-Noise Ratio (PSNR). Furthermore, HCSTNet demonstrates enhanced computational efficiency. It utilizes fewer parameters and requires less inference time than current state-of-the-art models. These findings validate HCSTNet as a highly effective and resource-efficient tool for channel estimation in real- world VLC applications. 본 연구는 OIRS(Optical Intelligent Reflective Surfaces)로 증강된 실내 MIMO 가시광통신(VLC) 네트워크에서 채널 상태를 추정하는 고급 방법을 조사합니다. VLC는 상당한 대역폭과 전자기 간섭에 대한 내성을 제공하지만, 정확한 채널 정보를 얻는 것은 어렵습니다. 이러한 어려움은 다중 경로 신호 반사, 비가시선(NLOS) 경로 및 복잡한 잡음 패턴에서 비롯됩니다. 이러한 문제를 극복하기 위해 본 연구에서는 새로운 하이브리드 딥러닝 아키텍처인 HCSTNet을 제안합니다. 이 프레임워크는 잡음 추정 전용 합성곱 신경망(CNN)과 잡음 제거를 위해 설계된 Swin Transformer 모듈을 결합합니다. 이러한 구성 요소를 결합함으로써 시스템은 채널 데이터 내의 지역적 잡음 세부 정보와 장거리 공간 상관 관계를 효과적으로 포착합니다. 본 연구에서는 가시선(LOS), 비가시선 및 OIRS 반사 경로를 고려한 실내 MIMO VLC 설정을 모델링하는 포괄적인 시뮬레이션 환경을 구축했습니다. 훈련 데이터에는 열 잡음 및 샷 잡음과 같은 현실적인 장애 요인이 포함됩니다. 실험 결과는 HCSTNet이 기존의 전통적인 방식 및 딥러닝 방식보다 우수한 성능을 보임을 나타냅니다. 특히, 이 모델은 정규화 평균 제곱 오차(NMSE)를 10⁻³~10⁻⁴ 범위로 줄이는 동시에 최대 신호 대 잡음비(PSNR)를 크게 향상시킵니다. 또한, HCSTNet은 향상된 계산 효율성을 보여줍니다. 현재 최첨단 모델보다 더 적은 파라미터를 사용하고 추론 시간도 단축됩니다. 이러한 결과는 HCSTNet이 실제 VLC(가시광 통신) 애플리케이션에서 채널 추정을 위한 매우 효과적이고 자원 효율적인 도구임을 입증합니다.

    • Semiconductor Defect Detection with Swin Transformer-Based Hybrid RCNN under Limited Labeled Data

      ERGASHOV ELDOR 아주대학교 정보통신대학원 2026 국내석사

      RANK : 232268

      Automated visual inspection is essential for maintaining quality control in semiconductor manufacturing, where even minor defects can lead to device failures and significant financial losses. This thesis presents a hybrid deep learning approach for semiconductor defect detection that combines the Swin Transformer backbone with the Cascade R-CNN detection framework. The proposed method addresses the challenge of detecting fine-grained structural defects in transistor components under limited labeled data conditions. The proposed Swin-T + Cascade R-CNN architecture leverages the hierarchical attention mechanism of the Swin Transformer to capture both local defect patterns and global structural context, while the multi-stage Cascade R-CNN detection head provides progressive bounding box refinement for precise defect localization. Transfer learning from ImageNet-pretrained weights and targeted data augmentation strategies enable effective training with only 172 labeled images. Experimental evaluation on the MVTec-AD transistor dataset demonstrates that the proposed method achieves a mAP@0.5:0.95 of 0.994, outperforming Faster R-CNN (0.975) and Cascade R-CNN with ResNet-50 backbone (0.989). The proposed method attains the highest recall (99.7%) and F1-score (97.2%) among evaluated models, with a defect classification accuracy of 96% on 100 test images. Per-class analysis shows significant improvements in detecting the defect class obj_head_ng (AP: 0.979) and challenging structural components. Robustness evaluation under Gaussian blur conditions indicates that the model maintains F1 scores above 0.96 even under moderate image degradation. These results validate the effectiveness of combining transformer-based feature extraction with cascade detection for industrial defect detection tasks. The proposed approach offers a practical solution for automated visual inspection in semiconductor manufacturing environments where detection accuracy and reliability are paramount.

    • Swin Transformer 및 ConvGRU 기반 다중작업 학습을 활용한 저조도 비디오 위장 객체 탐지 및 영상 복원 최적화 연구

      유석재 건국대학교 대학원 2026 국내석사

      RANK : 232239

      저조도 비디오 위장 객체 탐지는 저조도 환경에서 발생하는 노이즈, 정보 소실과 같은 시각적 열화와 위장 객체 고유의 모호성이 결합된 도전적인 연구 문제이다. 기존의 위장 객체 탐지 모델들은 고품질의 주간 비디오 입력을 전제로 설계되며, 객체의 명확한 텍스처와 경계 정보에 의존하기 때문에저조도 환경에서는 성능 저하를 보인다. 이는 저조도 감시, 악천후 등 저조도 조건에서의 객체 탐지 기술의 실제적 적용을 제한하는 한계로 작용한다. 본 연구는 이러한 한계를 극복하기 위해 Swin Transformer와 ConvGRU를 기반으로 한 ‘분해-인식 다중작업 학습’ 구조를 채택하였다. 해당 구조는 저조도 강화를 보조 작업으로, 위장 객체 분할을 주요 작업으로 설정한다. 보조 작업을 통해 모델이 장면의 소실된 구조 및 질감 정보를 복원하도록 강제함으로써, 두 작업 간의 공유 특징 표현을 견고하게 학습하도록 유도하는 것이 핵심이다. 또한, 본 연구는 다중작업 학습을 위한 현실적인 합성 데이터 생성 방식을 제안한다. 주요 실험을 위해 구축된 합성 저조도 데이터셋(LLCA-Mask) 및 기존 주간 데이터셋(MoCA-Mask, CAD)을 사용하여 제안 모델의 성능을 기존 모델들과 비교 평가하였다. 합성 저조도 데이터셋을 활용한 평가에서 기존 모델들은 위장 객체를 탐지하지 못하거나 노이즈를 객체로 오인, 이로 인해 S-measure가 최대 0.126 하락하는 등 저조도 환경에서 정확도가 하락하였다. 그에 비해 본 연구에서 제안하는 모델은 저조도 환경에서 기존 SOTA 모델보다 S-measure가 최대 0.139 (평균 0.0942) 상승하는 등 객체의 경계를 비교적 정확하게 분할해냈다. 본 연구는 저조도 비디오 위장 객체 분할 문제에 대한 복원과 분할의 동시 최적화라는 새로운 접근 방식을 제시하며, 향후 관련 연구를 위한 새로운 비교 척도를 제공한다는 점에서 의의가 있으며, 이는 열악한 저조도 환경에서의 강건한 위장 객체 탐지 및 분할 모델 개발에 기여할 것으로 기대된다.

    연관 검색어 추천

    이 검색어로 많이 본 자료

    활용도 높은 자료

    해외이동버튼