RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Refining Embeddings in Transformers : Image Classification, MRI Reconstruction, and Safe Image Generation = 트랜스포머 임베딩 정제 : 이미지 분류, MRI 재구성 및 안전한 이미지 생성

    한글로보기

    https://www.riss.kr/link?id=T17392903

    • 저자
    • 발행사항

      대구 : 경북대학교 대학원, 2026

    • 학위논문사항

      Thesis (doctoral) -- 경북대학교 대학원 , 인공지능학과 , 2026. 2

    • 발행연도

      2026

    • 작성언어

      영어

    • 주제어
    • DDC

      006.693 판사항(23)

    • 발행국(도시)

      대한민국

    • 기타서명

      트랜스포머 임베딩 정제 : 이미지 분류, MRI 재구성 및 안전한 이미지 생성

    • 형태사항

      x, 174 p. : ill., charts ; 26 cm.

    • 일반주기명

      Thesis Advisor: 정희철.
      Includes bibliographical references.

    • UCI식별코드

      I804:22001-000000111558

    • 소장기관
      • 경북대학교 중앙도서관 소장기관정보
    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Vision transformer research has evolved from foundational architectural designs to specialized domain applications and, more recently, to the critical issue of model safety. Following this trajectory, this thesis pursues three complementary lines of research to advance both the capability and responsible deployment of transformer models. First, to improve general-purpose vision backbones, I propose novel embedding architectures that utilize nonlinear transformations and shared-weight structures. Complementing these embedding-level improvements, a self-attention-based classifier head is introduced to replace the standard linear head, thereby improving feature aggregation at the final output stage. Second, adapting transformer architectures to the medical domain, I develop a domain-aware model for MRI reconstruction. This approach introduces a gated dual-domain transformer that processes spatial and frequency information concurrently to effectively correct off-resonance artifacts. Third, recognizing the safety risks in generative AI, I present a practical machine unlearning framework that controls embedding space. This method neutralizes unsafe concepts by realigning embeddings within the text encoder, ensuring robust defense against adversarial attacks while preserving benign generation quality. Collectively, these contributions demonstrate a comprehensive advancement of vision transformers, spanning from architectural refinement to domain specialization and safety alignment.
    번역하기

    Vision transformer research has evolved from foundational architectural designs to specialized domain applications and, more recently, to the critical issue of model safety. Following this trajectory, this thesis pursues three complementary lines of r...

    Vision transformer research has evolved from foundational architectural designs to specialized domain applications and, more recently, to the critical issue of model safety. Following this trajectory, this thesis pursues three complementary lines of research to advance both the capability and responsible deployment of transformer models. First, to improve general-purpose vision backbones, I propose novel embedding architectures that utilize nonlinear transformations and shared-weight structures. Complementing these embedding-level improvements, a self-attention-based classifier head is introduced to replace the standard linear head, thereby improving feature aggregation at the final output stage. Second, adapting transformer architectures to the medical domain, I develop a domain-aware model for MRI reconstruction. This approach introduces a gated dual-domain transformer that processes spatial and frequency information concurrently to effectively correct off-resonance artifacts. Third, recognizing the safety risks in generative AI, I present a practical machine unlearning framework that controls embedding space. This method neutralizes unsafe concepts by realigning embeddings within the text encoder, ensuring robust defense against adversarial attacks while preserving benign generation quality. Collectively, these contributions demonstrate a comprehensive advancement of vision transformers, spanning from architectural refinement to domain specialization and safety alignment.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    비전 트랜스포머 연구는 기초적인 아키텍처 설계에서 시작하여 특정 도메인으로의 응용, 그리고 최근에는 모델의 안전성 문제로 그 초점이 확장되어 왔다. 본 논문에서는 이러한 연구 흐름에 발맞추어 비전 트랜스포머의 성능과 활용성을 종합적으로 발전시키기 위한 세 가지 연구를 수행하였다. 첫째, 범용 비전 모델의 기초 역량을 강화하기 위해 임베딩 레이어와 분류기 구조를 개선하였다. 비선형 변환과 가중치 공유 구조를 도입하여 임베딩의 표현력을 높였으며, 기존의 선형 분류기를 대체하는 자기 주의 기반의 분류기 헤드를 도입하여 최종 단계에서의 특징 해석 성능을 개선하였다. 둘째, 의료 영상 분야의 특수한 요구사항을 반영하여 MRI 재구성을 위한 도메인 특화 트랜스포머를 개발하였다. 제안하는 게이트 이중 도메인 트랜스포머는 공간과 주파수 도메인 정보를 동시에 처리하여 MRI의 비공명 아티팩트를 효과적으로 보정한다. 셋째, 생성형 AI 모델의 안전한 활용을 위해 실용적인 머신 언러닝 프레임워크인 임베딩 공간 왜곡 기법을 제안하였다. 이 방법은 텍스트 인코더 내에서 유해 임베딩을 재정렬하여 생성 품질을 유지하면서도 유해 개념을 효과적으로 무력화한다. 결과적으로 본 논문은 아키텍처의 정교화부터 도메인 특화, 그리고 안전성 확보에 이르기까지 비전 트랜스포머 연구의 핵심적인 발전 방향을 포괄적으로 제시한다.
    번역하기

    비전 트랜스포머 연구는 기초적인 아키텍처 설계에서 시작하여 특정 도메인으로의 응용, 그리고 최근에는 모델의 안전성 문제로 그 초점이 확장되어 왔다. 본 논문에서는 이러한 연구 흐름...

    비전 트랜스포머 연구는 기초적인 아키텍처 설계에서 시작하여 특정 도메인으로의 응용, 그리고 최근에는 모델의 안전성 문제로 그 초점이 확장되어 왔다. 본 논문에서는 이러한 연구 흐름에 발맞추어 비전 트랜스포머의 성능과 활용성을 종합적으로 발전시키기 위한 세 가지 연구를 수행하였다. 첫째, 범용 비전 모델의 기초 역량을 강화하기 위해 임베딩 레이어와 분류기 구조를 개선하였다. 비선형 변환과 가중치 공유 구조를 도입하여 임베딩의 표현력을 높였으며, 기존의 선형 분류기를 대체하는 자기 주의 기반의 분류기 헤드를 도입하여 최종 단계에서의 특징 해석 성능을 개선하였다. 둘째, 의료 영상 분야의 특수한 요구사항을 반영하여 MRI 재구성을 위한 도메인 특화 트랜스포머를 개발하였다. 제안하는 게이트 이중 도메인 트랜스포머는 공간과 주파수 도메인 정보를 동시에 처리하여 MRI의 비공명 아티팩트를 효과적으로 보정한다. 셋째, 생성형 AI 모델의 안전한 활용을 위해 실용적인 머신 언러닝 프레임워크인 임베딩 공간 왜곡 기법을 제안하였다. 이 방법은 텍스트 인코더 내에서 유해 임베딩을 재정렬하여 생성 품질을 유지하면서도 유해 개념을 효과적으로 무력화한다. 결과적으로 본 논문은 아키텍처의 정교화부터 도메인 특화, 그리고 안전성 확보에 이르기까지 비전 트랜스포머 연구의 핵심적인 발전 방향을 포괄적으로 제시한다.

    더보기

    목차 (Table of Contents)

    • I. Introduction 1
    • II. Background 7
    • 2.1 Advancements in Vision Transformer Architectures 7
    • 2.1.1 Integrating Inductive Biases 7
    • 2.1.2 Addressing Quadratic Complexity 7
    • I. Introduction 1
    • II. Background 7
    • 2.1 Advancements in Vision Transformer Architectures 7
    • 2.1.1 Integrating Inductive Biases 7
    • 2.1.2 Addressing Quadratic Complexity 7
    • 2.2 Applications of Vision Transformers in the Medical Domain 9
    • 2.2.1 Application to Medical Image Analysis and Synthesis 9
    • 2.2.2 Application to MRI Reconstruction 9
    • 2.3 Safety Concerns of ViT-based Image Generation Models 11
    • 2.3.1 Adversarial Attacks 11
    • 2.3.2 Defense Methods 12
    • III.Advanced Architectural Components for Vision Transformers: Q, K, V Embeddings and Classifier Heads 15
    • 3.1 Motivation and Problem Formulation 15
    • 3.2 Nonlinear and Shared Embedding Architectures 16
    • 3.2.1 Self-Attention Mechanisms in Vision Transformers 17
    • 3.2.2 Proposed Method 19
    • 3.2.3 Experiments 24
    • 3.2.4 Analysis and Ablation Study 35
    • 3.2.5 Limitations 41
    • 3.3 Self-Attention Classifier Head 43
    • 3.3.1 Proposed Method 43
    • 3.3.2 Experiments 48
    • 3.4 Summary 57
    • IV. A Dual-Domain Vision Transformer for MRI Reconstruction 58
    • 4.1 Motivation and Problem Formulation 59
    • 4.2 Proposed Methods 61
    • 4.2.1 Dual Domain-aware Query, Key, and Value Embedding 61
    • 4.2.2 Selective Perceptual Loss 66
    • 4.2.3 Test-time Translation Merger 69
    • 4.3 Experiments 71
    • 4.3.1 Experimental Details 71
    • 4.3.2 Performance Evaluation on Simulated Off-resonance Dataset 76
    • 4.3.3 Performance Evaluation on Real Off-resonance Dataset 79
    • 4.4 Analysis 82
    • 4.4.1 Generalization to Unseen Off-resonance Frequencies 82
    • 4.4.2 Analysis of Effective Receptive Field 84
    • 4.4.3 Computational Costs 85
    • 4.5 Ablation Studies 87
    • 4.5.1 Encoder vs. Decoder 87
    • 4.5.2 Combinations of Input and Output 88
    • 4.5.3 Loss Functions 89
    • 4.5.4 Impact of VGG16 Feature Layer Selection 91
    • 4.5.5 What is the Role of α in LS? 93
    • 4.5.6 What is the Role of TTM? 94
    • 4.6 Discussion 100
    • 4.6.1 Validity of Natural Image-based Metrics for MRI 100
    • 4.6.2 Limitations 101
    • 4.7 Summary 103
    • V. Concept Unlearning in Transformer Embedding Space for Generative AI Safety 105
    • 5.1 Motivation and Problem Formulation 105
    • 5.2 Proposed Methods 107
    • 5.2.1 Target Vector Generation Phase 108
    • 5.2.2 Training Phase 111
    • 5.3 Experiments 115
    • 5.3.1 Experimental Settings 115
    • 5.3.2 Experimental Results on T2I 118
    • 5.3.3 Experimental Results on I2I 127
    • 5.3.4 Experimental Results on Other Concepts 128
    • 5.3.5 Generalizability across Diffusion Models 132
    • 5.3.6 Efficiency Evaluation 133
    • 5.3.7 Comprehensive Evaluation Results 134
    • 5.3.8 Combining with Basic Defense Methods 135
    • 5.3.9 Evaluation on Benign Images using Current Assessment Methods 136
    • 5.4 Further Analysis 137
    • 5.4.1 Analysis of Embedding Space Distortion 137
    • 5.4.2 Analysis of Target Safe Vector Generation Phase 139
    • 5.4.3 CLIP Score and FID for Prompts Closer to “Nudity” Embedding 140
    • 5.4.4 Failure Case Analysis 142
    • 5.5 Ablation Studies 144
    • 5.5.1 Contributions of Each Loss Function 144
    • 5.5.2 Effect of the Loss Adjustment Technique 144
    • 5.5.3 Effect of Loss Coefficient λ 145
    • 5.5.4 Effective Target of λ 146
    • 5.6 Remarks 147
    • 5.6.1 Cultural Bias 147
    • 5.6.2 Limitations 148
    • 5.6.3 Broader Impacts 149
    • 5.7 Summary 149
    • VI. Conclusion and Future Works 151
    • References 155
    • Abstract (In English) 171
    • Abstract (In Korean) 173
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼