RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    ALOHA: Anomaly-generating Latent Operator with Hierarchical Attention for Fine-grained Defect Synthesis = 산업 미세 결함 이미지 생성을 위한 계층적 어텐션 기반 잠재 확산 모델

    한글로보기

    https://www.riss.kr/link?id=T17452037

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    AI-powered visual inspection in industrial manufacturing suffers from the scarcity and subtleness of defect samples. Consequently, anomaly generation has become crucial for building robust inspection systems, yet it remains challenging when dealing with fine-grained anomalies such as tiny, slender defects, which increasingly emerge as manufacturing processes continue to miniaturize. Despite recent progress in diffusion-based generators, it remains difficult to synthesize or preserve such fine structures, as standard latent diffusion models often sacrifice high-frequency spatial details during the encoding process.
    In this work, we present a novel diffusion operator, Anomaly-generating Latent Operator with Hierarchical Attention (ALOHA), designed to preserve such fine-grained anomalies. Our key insight is to leverage Cross-scale Fusion using Receptive-Field–augmented Attention (RFAtten) to learn richly fine-grained anomaly information across all features coming from the hierarchical U-Net denoiser. This mechanism effectively bridges the information gap between shallow and deep layers, preventing the dilution of minute structural cues during the denoising steps. By dynamically calibrating the receptive fields, RFAtten enables the model to simultaneously capture fine-grained morphological structures while maintaining global semantic consistency. ALOHA integrates a ControlNet branch for mask-guided spatial conditioning to enable the generation of accurately-matched anomalous image-mask pairs and to achieve precise, mask-guided spatial control.
    This synergistic architecture ensures that the synthesized defects not only blend seamlessly with the background texture but also strictly adhere to the morphological constraints of the input masks. Extensive experiments on the MVTec AD dataset demonstrate that ALOHA generates highly authentic and structurally coherent anomalies, outperforming state-of-the-art methods and significantly improving downstream anomaly inspection performance.
    번역하기

    AI-powered visual inspection in industrial manufacturing suffers from the scarcity and subtleness of defect samples. Consequently, anomaly generation has become crucial for building robust inspection systems, yet it remains challenging when dealing wi...

    AI-powered visual inspection in industrial manufacturing suffers from the scarcity and subtleness of defect samples. Consequently, anomaly generation has become crucial for building robust inspection systems, yet it remains challenging when dealing with fine-grained anomalies such as tiny, slender defects, which increasingly emerge as manufacturing processes continue to miniaturize. Despite recent progress in diffusion-based generators, it remains difficult to synthesize or preserve such fine structures, as standard latent diffusion models often sacrifice high-frequency spatial details during the encoding process.
    In this work, we present a novel diffusion operator, Anomaly-generating Latent Operator with Hierarchical Attention (ALOHA), designed to preserve such fine-grained anomalies. Our key insight is to leverage Cross-scale Fusion using Receptive-Field–augmented Attention (RFAtten) to learn richly fine-grained anomaly information across all features coming from the hierarchical U-Net denoiser. This mechanism effectively bridges the information gap between shallow and deep layers, preventing the dilution of minute structural cues during the denoising steps. By dynamically calibrating the receptive fields, RFAtten enables the model to simultaneously capture fine-grained morphological structures while maintaining global semantic consistency. ALOHA integrates a ControlNet branch for mask-guided spatial conditioning to enable the generation of accurately-matched anomalous image-mask pairs and to achieve precise, mask-guided spatial control.
    This synergistic architecture ensures that the synthesized defects not only blend seamlessly with the background texture but also strictly adhere to the morphological constraints of the input masks. Extensive experiments on the MVTec AD dataset demonstrate that ALOHA generates highly authentic and structurally coherent anomalies, outperforming state-of-the-art methods and significantly improving downstream anomaly inspection performance.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    산업 제조 공정에서의 AI 기반 시각 검사 시스템은 결함 샘플의 부족성과 결함 자체의 미세한 특성으로 인해 성능 향상에 한계를 겪고 있다. 이에 따라 이상 이미지 생성은 견고한 검사 시스템 구축을 위한 핵심 요소로 부상하였으나, 공정의 소형화로 인해 점점 더 빈번하게 발생하는 미세한 결함을 효과적으로 생성하는 것은 여전히 도전적인 문제로 남아 있다. 특히 기존 잠재 확산 모델은 인코딩 과정에서 고주파 공간 정보를 손실하는 경향이 있어 이러한 미세 구조를 합성하거나 보존하는 데 한계를 보인다.
    본 논문에서는 미세 결함 정보를 효과적으로 보존하기 위한 새로운 잠재 확산 기반 모델인 계층적 어텐션 기반 이상 생성 잠재 연산자를 제안한다. 제안한 방법은 수용영역 확장 어텐션을 활용한 교차 스케일 융합을 통해 계층적 U-Net의 얕은 계층과 깊은 계층 간 정보 단절을 완화하고, 확산 복원 과정에서 미세 구조 정보가 희석되는 현상을 효과적으로 방지한다. 또한 수용영역을 동적으로 조절함으로써 미세한 형태적 구조를 정밀하게 포착함과 동시에 전역적인 의미적 일관성을 유지한다.
    더 나아가 계층적 어텐션 기반 이상 생성 잠재 연산자는 마스크 기반 공간 조건화를 위해 ControlNet 분기를 통합하여, 결함 영역이 정확히 제어된 이상 이미지–마스크 쌍을 생성할 수 있도록 설계되었다. 이를 통해 생성된 결함은 배경 질감과 자연스럽게 융합되는 동시에 입력 마스크에 정의된 형태적 제약을 엄격하게 만족한다. MVTec AD 데이터셋을 대상으로 한 실험 결과, 제안한 방법은 기존 최신 이상 생성 기법들 대비 더욱 사실적이고 구조적으로 일관된 이상 이미지를 생성하며, 후속 이상 탐지 성능을 유의미하게 향상시킴을 확인하였다.
    번역하기

    산업 제조 공정에서의 AI 기반 시각 검사 시스템은 결함 샘플의 부족성과 결함 자체의 미세한 특성으로 인해 성능 향상에 한계를 겪고 있다. 이에 따라 이상 이미지 생성은 견고한 검사 시스...

    산업 제조 공정에서의 AI 기반 시각 검사 시스템은 결함 샘플의 부족성과 결함 자체의 미세한 특성으로 인해 성능 향상에 한계를 겪고 있다. 이에 따라 이상 이미지 생성은 견고한 검사 시스템 구축을 위한 핵심 요소로 부상하였으나, 공정의 소형화로 인해 점점 더 빈번하게 발생하는 미세한 결함을 효과적으로 생성하는 것은 여전히 도전적인 문제로 남아 있다. 특히 기존 잠재 확산 모델은 인코딩 과정에서 고주파 공간 정보를 손실하는 경향이 있어 이러한 미세 구조를 합성하거나 보존하는 데 한계를 보인다.
    본 논문에서는 미세 결함 정보를 효과적으로 보존하기 위한 새로운 잠재 확산 기반 모델인 계층적 어텐션 기반 이상 생성 잠재 연산자를 제안한다. 제안한 방법은 수용영역 확장 어텐션을 활용한 교차 스케일 융합을 통해 계층적 U-Net의 얕은 계층과 깊은 계층 간 정보 단절을 완화하고, 확산 복원 과정에서 미세 구조 정보가 희석되는 현상을 효과적으로 방지한다. 또한 수용영역을 동적으로 조절함으로써 미세한 형태적 구조를 정밀하게 포착함과 동시에 전역적인 의미적 일관성을 유지한다.
    더 나아가 계층적 어텐션 기반 이상 생성 잠재 연산자는 마스크 기반 공간 조건화를 위해 ControlNet 분기를 통합하여, 결함 영역이 정확히 제어된 이상 이미지–마스크 쌍을 생성할 수 있도록 설계되었다. 이를 통해 생성된 결함은 배경 질감과 자연스럽게 융합되는 동시에 입력 마스크에 정의된 형태적 제약을 엄격하게 만족한다. MVTec AD 데이터셋을 대상으로 한 실험 결과, 제안한 방법은 기존 최신 이상 생성 기법들 대비 더욱 사실적이고 구조적으로 일관된 이상 이미지를 생성하며, 후속 이상 탐지 성능을 유의미하게 향상시킴을 확인하였다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Table of Contents ii
    • List of Tables iv
    • List of Figures v
    • Chapter 1. Introduction 1
    • Abstract i
    • Table of Contents ii
    • List of Tables iv
    • List of Figures v
    • Chapter 1. Introduction 1
    • 1.1. Background 1
    • 1.2. Motivation of the Dissertation 6
    • 1.3. Research Outline 9
    • Chapter 2. Related work 11
    • 2.1. Traditional Anomaly Generation Methods 11
    • 2.2. GAN-based Anomaly Generation Methods 13
    • 2.3. Diffusion-based Anomaly Generation Methods 15
    • Chapter 3. Method 17
    • 3.1. Preliminary 18
    • 3.1.1. Latent Diffusion Models 18
    • 3.1.2. ControlNet 20
    • 3.1.3. Vision Transformer 23
    • 3.2. Receptive-Field-augmented Attention 25
    • 3.2.1. Multi-Head Self-Attention with Convolutions 26
    • 3.2.2. Receptive Field Block (RFB) 28
    • 3.3. Cross-scale Fusion in U-Net denoiser 30
    • 3.4. Learning Anomaly 33
    • 3.4.1. Setup 33
    • 3.4.2. Conditioning 33
    • 3.4.3. Objective 33
    • 3.4.4. Training scheme 34
    • Chapter 4. Experiments 35
    • 4.1. Experiment Settings 35
    • 4.1.1. Dataset 35
    • 4.1.2. Implementation Details 35
    • 4.1.3. Metric 35
    • 4.1.4. Baselines 36
    • 4.2. Anomaly Generation Evaluation 37
    • 4.2.1. Qualitative Results 37
    • 4.2.2. Quantitative Results 40
    • 4.3. Anomaly Inpection Evaluation 42
    • 4.3.1. Anomaly Localization Result 42
    • 4.3.2. Anomaly Classification Result 42
    • 4.4. Additional Qualitative Results 44
    • 4.4.1. MVTec AD Dataset 44
    • 4.4.2. VisA Dataset 44
    • 4.4.3. Comparison on 80%-Scaled Masks 45
    • 4.5. Additional Quantitative Results 58
    • 4.5.1. Anomaly Detection 58
    • 4.6. Ablation Study 59
    • 4.6.1. Attention Module 59
    • 4.6.2. Cross-scale Fusion 61
    • 4.6.3. Fusion Strategy 62
    • Chapter 5. Conclusion 65
    • 5.1. Conclusion 65
    • 5.2. Limitation 67
    • Bibliography 69
    • 국문 초록 72
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼