RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Context-preserving Concept Erasure in Diffusion Models = 이미지 생성 모델에서 주변 맥락을 보존하는 개념 제거 기법 연구

    한글로보기

    https://www.riss.kr/link?id=T17449972

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    대규모 텍스트-이미지(Text-to-Image, T2I) 확산(diffusion) 모델은 자연어 프롬프트로부터 고품질의 다양한 이미지를 생성하는 데 성공을 거두며, 예술·디자인·엔터테인먼트 등 여러 창작 분야에서 폭넓게 활용되고 있다. 그러나 이러한 모델의 개방성과 접근성은 저작권 침해, 인물 이미지의 오용, 부적절하거나 유해한 콘텐츠 생성 등 심각한 윤리적·법적 문제를 야기하고 있다. 이러한 문제를 완화하기 위해 최근 연구에서는 사전 학습된 확산 모델로부터 특정 개념(예: 인물, 예술적 스타일, 노출 또는 폭력적 요소 등)을 제거하면서도 생성 품질을 유지하려는 개념 제거(concept erasure) 기법이 주목받고 있다.

    기존의 대부분 접근법은 텍스트 임베딩과 이미지 특징을 정렬하는 cross-attention 계층만을 미세 조정하여, 관련 없는 개념에 대한 간섭을 줄이는 방식을 사용한다. 그러나 이러한 국소적 업데이트만으로는 목표 개념을 완전히 분리하기 어렵고, 종종 두 가지 부작용을 초래한다. 첫째, 의미적 변형(semantic drift)은 비관련 텍스트 표현의 전역적 의미가 변화하는 현상이며, 둘째, 공간적 왜곡(spatial distortion)은 이미지 내의 관계없는 영역이 의도치 않게 변형되는 문제를 의미한다.

    본 논문은 이러한 한계를 해결하기 위해 주변 문맥을 보존하는 개념 제거(context-preserving concept erasure)를 의미적(semantic) 관점과 공간적(spatial) 관점에서 상보적으로 탐구한다.
    의미적 수준에서는 cross-attention 계층 내에 Residual Attention Gate (ResAG)를 도입하여 목표 개념과 관련된 활성화를 선택적으로 조절하고, Robust Adversarial Residual Erasure (RARE) 모듈을 통해 제거된 개념이 프롬프트 기반의 적대적 공격으로 재생성되는 것을 방지한다. RARE는 잔존 임베딩 방향을 반복적으로 탐색·억제함으로써 높은 강인성을 확보한다. 공간적 수준에서는 학습이 필요 없는 Gated Low-rank Concept Erasure (GLoCE) 기법을 제안한다. 이는 diffusion 과정 중 특정 층의 임베딩이 목표 개념의 주성분(principal components)과 정렬될 때 low-rank adaptor를 활성화하여, 목표 영역만 국소적으로 억제하면서 주변의 구조적·시각적 일관성을 유지한다.

    이 두 접근법—의미적 제어를 수행하는 CPE와 공간적 제거를 수행하는 GLoCE—는 모두 “정확하고(precise), 강인하며(robust), 문맥을 보존하는(context-preserving)” 개념 제거를 지향한다.
    본 연구는 비선형 게이팅과 저차원 적응 기법을 결합함으로써, 확산 모델 내 개념 분리(disentanglement)와 제어 가능성(controllability)에 대한 이해를 심화시키고, 보다 안전하고 해석 가능한 생성형 인공지능 시스템 구축의 기반을 제시한다.
    번역하기

    대규모 텍스트-이미지(Text-to-Image, T2I) 확산(diffusion) 모델은 자연어 프롬프트로부터 고품질의 다양한 이미지를 생성하는 데 성공을 거두며, 예술·디자인·엔터테인먼트 등 여러 창작 분야에�...

    대규모 텍스트-이미지(Text-to-Image, T2I) 확산(diffusion) 모델은 자연어 프롬프트로부터 고품질의 다양한 이미지를 생성하는 데 성공을 거두며, 예술·디자인·엔터테인먼트 등 여러 창작 분야에서 폭넓게 활용되고 있다. 그러나 이러한 모델의 개방성과 접근성은 저작권 침해, 인물 이미지의 오용, 부적절하거나 유해한 콘텐츠 생성 등 심각한 윤리적·법적 문제를 야기하고 있다. 이러한 문제를 완화하기 위해 최근 연구에서는 사전 학습된 확산 모델로부터 특정 개념(예: 인물, 예술적 스타일, 노출 또는 폭력적 요소 등)을 제거하면서도 생성 품질을 유지하려는 개념 제거(concept erasure) 기법이 주목받고 있다.

    기존의 대부분 접근법은 텍스트 임베딩과 이미지 특징을 정렬하는 cross-attention 계층만을 미세 조정하여, 관련 없는 개념에 대한 간섭을 줄이는 방식을 사용한다. 그러나 이러한 국소적 업데이트만으로는 목표 개념을 완전히 분리하기 어렵고, 종종 두 가지 부작용을 초래한다. 첫째, 의미적 변형(semantic drift)은 비관련 텍스트 표현의 전역적 의미가 변화하는 현상이며, 둘째, 공간적 왜곡(spatial distortion)은 이미지 내의 관계없는 영역이 의도치 않게 변형되는 문제를 의미한다.

    본 논문은 이러한 한계를 해결하기 위해 주변 문맥을 보존하는 개념 제거(context-preserving concept erasure)를 의미적(semantic) 관점과 공간적(spatial) 관점에서 상보적으로 탐구한다.
    의미적 수준에서는 cross-attention 계층 내에 Residual Attention Gate (ResAG)를 도입하여 목표 개념과 관련된 활성화를 선택적으로 조절하고, Robust Adversarial Residual Erasure (RARE) 모듈을 통해 제거된 개념이 프롬프트 기반의 적대적 공격으로 재생성되는 것을 방지한다. RARE는 잔존 임베딩 방향을 반복적으로 탐색·억제함으로써 높은 강인성을 확보한다. 공간적 수준에서는 학습이 필요 없는 Gated Low-rank Concept Erasure (GLoCE) 기법을 제안한다. 이는 diffusion 과정 중 특정 층의 임베딩이 목표 개념의 주성분(principal components)과 정렬될 때 low-rank adaptor를 활성화하여, 목표 영역만 국소적으로 억제하면서 주변의 구조적·시각적 일관성을 유지한다.

    이 두 접근법—의미적 제어를 수행하는 CPE와 공간적 제거를 수행하는 GLoCE—는 모두 “정확하고(precise), 강인하며(robust), 문맥을 보존하는(context-preserving)” 개념 제거를 지향한다.
    본 연구는 비선형 게이팅과 저차원 적응 기법을 결합함으로써, 확산 모델 내 개념 분리(disentanglement)와 제어 가능성(controllability)에 대한 이해를 심화시키고, 보다 안전하고 해석 가능한 생성형 인공지능 시스템 구축의 기반을 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large-scale text-to-image (T2I) diffusion models have achieved remarkable success in generating high-quality, diverse images from natural language prompts. Their widespread adoption across creative domains—such as art, design, and entertainment—highlights their transformative potential.
    However, their open accessibility has also raised critical ethical and legal concerns, including copyright infringement, identity misuse, and the generation of explicit or harmful content. To mitigate such issues, recent research has explored concept erasure, which aims to remove specific concepts from pretrained diffusion models while preserving their overall generative quality.

    Most existing approaches constrain fine-tuning to the cross-attention layers that align text embeddings with image features, thereby reducing interference with unrelated semantics. However, such localized updates remain insufficient to completely isolate target concepts and often lead to two major side effects: semantic drift—global alteration of unrelated text representations—and spatial distortion—unintended modification of irrelevant image regions.

    To address these challenges, this thesis approaches context-preserving concept erasure from two complementary perspectives: semantic and spatial.
    At the semantic level, we introduce Residual Attention Gates (ResAGs) within the cross-attention layers, enabling selective modulation of target-related activations. A Robust Adversarial Residual Erasure (RARE) scheme further enhances robustness by iteratively detecting and suppressing residual embedding directions that can regenerate the erased concept, providing strong resistance against prompt-based adversarial attacks.

    At the spatial level, we develop a training-free Gated Low-rank Concept Erasure (GLoCE) mechanism that activates when layer embeddings align with the principal components of a target concept. This low-rank projection locally suppresses target-associated features while maintaining the structural and visual coherence of surrounding content.

    These two complementary approaches—semantic-level control via CPE and spatially-aware erasure via GLoCE—share a unified goal: to make concept removal in diffusion models precise, robust, and context-preserving.
    By integrating nonlinear gating and low-rank adaptation, this work advances the understanding of concept disentanglement and controllability in diffusion models, laying the groundwork for safer, more interpretable, and socially responsible generative systems.
    번역하기

    Large-scale text-to-image (T2I) diffusion models have achieved remarkable success in generating high-quality, diverse images from natural language prompts. Their widespread adoption across creative domains—such as art, design, and entertainment—hi...

    Large-scale text-to-image (T2I) diffusion models have achieved remarkable success in generating high-quality, diverse images from natural language prompts. Their widespread adoption across creative domains—such as art, design, and entertainment—highlights their transformative potential.
    However, their open accessibility has also raised critical ethical and legal concerns, including copyright infringement, identity misuse, and the generation of explicit or harmful content. To mitigate such issues, recent research has explored concept erasure, which aims to remove specific concepts from pretrained diffusion models while preserving their overall generative quality.

    Most existing approaches constrain fine-tuning to the cross-attention layers that align text embeddings with image features, thereby reducing interference with unrelated semantics. However, such localized updates remain insufficient to completely isolate target concepts and often lead to two major side effects: semantic drift—global alteration of unrelated text representations—and spatial distortion—unintended modification of irrelevant image regions.

    To address these challenges, this thesis approaches context-preserving concept erasure from two complementary perspectives: semantic and spatial.
    At the semantic level, we introduce Residual Attention Gates (ResAGs) within the cross-attention layers, enabling selective modulation of target-related activations. A Robust Adversarial Residual Erasure (RARE) scheme further enhances robustness by iteratively detecting and suppressing residual embedding directions that can regenerate the erased concept, providing strong resistance against prompt-based adversarial attacks.

    At the spatial level, we develop a training-free Gated Low-rank Concept Erasure (GLoCE) mechanism that activates when layer embeddings align with the principal components of a target concept. This low-rank projection locally suppresses target-associated features while maintaining the structural and visual coherence of surrounding content.

    These two complementary approaches—semantic-level control via CPE and spatially-aware erasure via GLoCE—share a unified goal: to make concept removal in diffusion models precise, robust, and context-preserving.
    By integrating nonlinear gating and low-rank adaptation, this work advances the understanding of concept disentanglement and controllability in diffusion models, laying the groundwork for safer, more interpretable, and socially responsible generative systems.

    더보기

    목차 (Table of Contents)

    • Abstract
    • Chapter 1. Introduction 1
    • 1.1 Motivation 1
    • 1.2 Key Contributions 3
    • Chapter 2. Background 6
    • Abstract
    • Chapter 1. Introduction 1
    • 1.1 Motivation 1
    • 1.2 Key Contributions 3
    • Chapter 2. Background 6
    • 2.1 Background on Diffusion Models 6
    • 2.1.1 Latent Diffusion and Stable Diffusion 6
    • 2.1.2 Text–Image Alignment via Cross-Attention 7
    • 2.1.3 Limitation of Fine-tuning CA Layers in Concept Erasure 7
    • 2.2 Related Work 9
    • 2.2.1 Safe T2I Image Generation 9
    • 2.2.2 Fine-tuning-based Concept Erasing 9
    • 2.2.3 Preserving Remaining Concepts 10
    • Chapter 3. Semantic-level concept erasure via residual attention gating 11
    • 3.1 Semantic-level Concept Erasure via Residual Attention Gate 11
    • 3.1.1 Residual Attention Gate (ResAG) 11
    • 3.1.2 Loss Design 13
    • 3.1.3 Robust Training for CPE 14
    • 3.1.4 Efficiency of Gating for Semantic Concept Erasure 15
    • 3.2 Experiments 17
    • 3.2.1 Evaluation Metrics 18
    • 3.2.2 Celebrities Erasure 18
    • 3.2.3 Artistic Styles Erasure 19
    • 3.2.4 Explicit Content Erasure 21
    • 3.2.5 Robust Erasure Against Adversarial Prompts 22
    • 3.2.6 Ablation Studies 22
    • 3.2.7 Efficiency Analysis 23
    • Chapter 4. Spatial-level concept erasure via gated low-rank adaptation 25
    • 4.1 Motivation for Spatial Gating 25
    • 4.2 Localized Concept Erasure 25
    • 4.2.1 Closed-form Low-Rank Adaptation 26
    • 4.3 Gating for Spatial Concept Erasure 27
    • 4.3.1 Gate Mechanism for Specificity 27
    • 4.3.2 Gate Mechanism via Principal Components 28
    • 4.3.3 Inference-OnlyUpdateofGate 29
    • 4.4 Empirical Analysis of Gate Parameters 30
    • 4.4.1 Few-shot Parameter Estimation 30
    • 4.4.2 Hyper-parameter Sensitivity of α, β, and γ 31
    • 4.4.3 Ranks of Low-Rank Matrices 32
    • 4.4.4 Number of Generations per Concept 33
    • 4.4.5 Range of Diffusion Timesteps 34
    • 4.5 Experiments 35
    • 4.5.1 Celebrities Erasure 37
    • 4.5.2 Explicit Contents Erasure 38
    • 4.5.3 Robustness against Adversarial Attacks 39
    • 4.5.4 Artistic Styles Erasure 39
    • 4.5.5 Ablation Studies 41
    • Chapter 5. Conclusion 43
    • 5.1 Summary 43
    • 5.2 Future Work 44
    • Appendix A. Appendix of Semantic Concept Eraser 45
    • A.1 Further Discussions on CPE 45
    • A.1.1 Sampling method for anchoring concepts 45
    • A.1.2 Efficiency Study 47
    • A.2 Additional Ablation Studies for CPE 50
    • A.3 Additional Qualitative Results 53
    • A.3.1 Celebrities Erasure 53
    • A.3.2 Artistic Styles Erasure 54
    • A.3.3 Robustness on Explicit Styles Erasure 55
    • Appendix B. Appendix of Spatial Concept Eraser 56
    • B.1 Further Discussions on Spatial Concept Erasure 56
    • B.1.1 Extension to Multiple Concepts Erasure 56
    • B.1.2 Selection of Concepts 56
    • B.1.3 Efficiency Study 62
    • B.2 Details of Inference-Only Update of GLoCE 63
    • B.3 Illustration of Gate Activation Map 66
    • B.4 Additional Qualitative Results on Localized Celebrities Erasure 67
    • B.4.1 Example 1 of Single Celebrity Erasure 67
    • B.4.2 Example 2 of Single Celebrity Erasure 68
    • B.5 Additional Qualitative Results on Explicit Concept Erasure 69
    • B.6 Additional Qualitative Results on Robustness 70
    • 요약 78
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼