RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Advancing Responsible AI via Adversarial Mechanisms: From Data-Aware Training to Inference-Time Correction = 적대적 기법을 통한 책임 있는 인공지능 발전: 데이터 인지 학습 및 추론 교정

    한글로보기

    https://www.riss.kr/link?id=T17449778

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    As Artificial Intelligence (AI) systems transition from academic curiosities to integral components of society, their deployment in safety-critical domains, such as autonomous driving and medical diagnostics, is rapidly accelerating. In these high-stakes environments, where errors can have profound, real-world consequences on human well-being and societal trust, ensuring model reliability and trustworthiness becomes a paramount objective. A cornerstone of this endeavor is adversarial robustness—the ability of a model to maintain correct and stable predictions even when faced with small, maliciously crafted perturbations to its input. While this property is essential for building dependable AI, the most common and foundational defense, Adversarial Training (AT), is far from a perfect solution. This dissertation argues that standard AT exhibits significant latent flaws, including compromises in fairness and robust generalization, which can undermine model safety in unexpected and subtle ways. This research directly confronts these challenges. It advances the field by first deeply investigating these critical flaws and then, in a novel pivot, repurposes the very principles of adversarial attacks as a constructive tool for model correction and enhancement.

    This research presents a three-part approach. First, in Chapter 3, the critical robust fairness problem is addressed, where standard adversarial training disproportionately harms the robustness of certain classes, creating dangerous blind spots. This dissertation introduces Distance-Aware Fair Adversarial training (DAFA), a novel methodology that makes the training process sensitive to the underlying inter-class similarity. DAFA dynamically assigns class-specific loss weights and adaptive adversarial margins based on the geometric relationships between classes in the feature space. This mechanism effectively improves worst-class robustness and promotes a more equitable defense without degrading average performance. Second, in Chapter 4, the counter-intuitive robust overfitting phenomenon is investigated. This occurs when models begin to memorize hard-to-learn examples during adversarial training, which paradoxically harms robust generalization on unseen data. This research introduces Difficulty Proportional Label Smoothing (DPLS), an adaptive regularization technique designed to combat this. DPLS mitigates memorization by applying a label smoothing factor directly proportional to the measured difficulty of each training example. This intervention forces the model to learn more generalizable, broadly applicable robust features rather than memorizing noisy specifics. Finally, in Chapter 5, this dissertation pivots from defensive training to constructive correction, addressing the pervasive challenge of object hallucination in Large Vision-Language Models (LVLMs). Adversarial Visual Contrastive Decoding (AVCD) is introduced, a novel, training-free inference methodology. AVCD ingeniously repurposes an adversarial attack not to fool the model, but to generate a hard negative visual input—one that maximally induces hallucination. By contrasting the logit distribution from this adversarial, hallucination-prone state against the original, clean-input logits, AVCD precisely suppresses visually ungrounded tokens and significantly enhances model faithfulness at inference time.

    Collectively, this dissertation demonstrates the profound versatility of adversarial principles. By first refining adversarial training with data-aware and adaptive techniques (DAFA, DPLS) to shore up its critical weaknesses, and then repurposing these principles for inference-time correction (AVCD), this work provides a robust toolkit. The contributions herein offer tangible pathways for building fairer, more generalizable, and ultimately more reliable AI systems poised for real-world deployment.
    번역하기

    As Artificial Intelligence (AI) systems transition from academic curiosities to integral components of society, their deployment in safety-critical domains, such as autonomous driving and medical diagnostics, is rapidly accelerating. In these high-sta...

    As Artificial Intelligence (AI) systems transition from academic curiosities to integral components of society, their deployment in safety-critical domains, such as autonomous driving and medical diagnostics, is rapidly accelerating. In these high-stakes environments, where errors can have profound, real-world consequences on human well-being and societal trust, ensuring model reliability and trustworthiness becomes a paramount objective. A cornerstone of this endeavor is adversarial robustness—the ability of a model to maintain correct and stable predictions even when faced with small, maliciously crafted perturbations to its input. While this property is essential for building dependable AI, the most common and foundational defense, Adversarial Training (AT), is far from a perfect solution. This dissertation argues that standard AT exhibits significant latent flaws, including compromises in fairness and robust generalization, which can undermine model safety in unexpected and subtle ways. This research directly confronts these challenges. It advances the field by first deeply investigating these critical flaws and then, in a novel pivot, repurposes the very principles of adversarial attacks as a constructive tool for model correction and enhancement.

    This research presents a three-part approach. First, in Chapter 3, the critical robust fairness problem is addressed, where standard adversarial training disproportionately harms the robustness of certain classes, creating dangerous blind spots. This dissertation introduces Distance-Aware Fair Adversarial training (DAFA), a novel methodology that makes the training process sensitive to the underlying inter-class similarity. DAFA dynamically assigns class-specific loss weights and adaptive adversarial margins based on the geometric relationships between classes in the feature space. This mechanism effectively improves worst-class robustness and promotes a more equitable defense without degrading average performance. Second, in Chapter 4, the counter-intuitive robust overfitting phenomenon is investigated. This occurs when models begin to memorize hard-to-learn examples during adversarial training, which paradoxically harms robust generalization on unseen data. This research introduces Difficulty Proportional Label Smoothing (DPLS), an adaptive regularization technique designed to combat this. DPLS mitigates memorization by applying a label smoothing factor directly proportional to the measured difficulty of each training example. This intervention forces the model to learn more generalizable, broadly applicable robust features rather than memorizing noisy specifics. Finally, in Chapter 5, this dissertation pivots from defensive training to constructive correction, addressing the pervasive challenge of object hallucination in Large Vision-Language Models (LVLMs). Adversarial Visual Contrastive Decoding (AVCD) is introduced, a novel, training-free inference methodology. AVCD ingeniously repurposes an adversarial attack not to fool the model, but to generate a hard negative visual input—one that maximally induces hallucination. By contrasting the logit distribution from this adversarial, hallucination-prone state against the original, clean-input logits, AVCD precisely suppresses visually ungrounded tokens and significantly enhances model faithfulness at inference time.

    Collectively, this dissertation demonstrates the profound versatility of adversarial principles. By first refining adversarial training with data-aware and adaptive techniques (DAFA, DPLS) to shore up its critical weaknesses, and then repurposing these principles for inference-time correction (AVCD), this work provides a robust toolkit. The contributions herein offer tangible pathways for building fairer, more generalizable, and ultimately more reliable AI systems poised for real-world deployment.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    인공지능(AI) 시스템이 학술적 호기심의 대상을 넘어 사회의 핵심 구성 요소로 자리 잡으면서, 자율주행이나 의료 진단과 같은 안전이 필수적인 도메인으로의 배치가 가속화되고 있다. 이러한 고위험 환경에서는 AI의 오류가 인간의 안녕과 사회적 신뢰에 심각한, 그리고 실질적인 영향을 미칠 수 있으므로, 모델의 신뢰성과 안정성을 확보하는 것은 타협 불가능한 최우선 목표가 된다. 이러한 노력의 핵심 기반 중 하나는 적대적 강건성으로, 이는 모델이 입력에 대해 작고 악의적으로 조작된 교란이 가해지는 상황에서도 정확하고 안정적인 예측을 유지하는 능력을 의미한다. 이 특성은 신뢰할 수 있는 AI를 구축하는 데 필수적이지만, 가장 보편적이고 기초적인 방어 기법인 적대적 학습은 완벽한 해결책과는 거리가 멀다. 본 학위논문은 표준적인 적대적 학습이 공정성 및 강건한 일반화 성능을 저해하는 등, 예기치 못한 미묘한 방식으로 모델의 안전성을 훼손할 수 있는 중대한 잠재적 결함들을 내포하고 있음을 주장한다. 본 연구는 이러한 난제들을 정면으로 다룬다. 이 중대한 결함들을 먼저 깊이 있게 조사하고, 나아가 적대적 공격의 핵심 원리를 모델의 교정 및 강화를 위한 건설적인 도구로 재활용하는 새로운 관점으로의 전환을 제시하며 이 분야의 발전에 기여하고자 한다.

    본 연구는 세 부분의 접근 방식을 제시한다. 첫째, 3장에서는 표준적인 적대적 학습이 특정 클래스의 강건성을 불균형적으로 저해하여 위험한 사각지대를 발생시키는 심각한 강건한 공정성 문제를 다룬다. 본 논문은 클래스 간의 유사성을 학습 과정에 민감하게 반영하는 새로운 방법론인 거리 인지 기반 공정 적대적 학습 (DAFA)을 제안한다. DAFA는 특성 공간 내 클래스 간의 기하학적 관계를 기반으로 클래스별 손실 가중치와 적응형 적대적 마진을 동적으로 할당한다. 이 메커니즘은 평균 성능의 저하 없이 취약 클래스의 강건성을 효과적으로 향상시켜, 보다 공평한 방어 모델을 구현한다.
    둘째, 4장에서는 모델이 적대적 학습 중 학습하기 어려운 예제들을 암기하기 시작하며, 이로 인해 역설적으로 보이지 않는 데이터에 대한 강건한 일반화 성능이 저해되는 직관에 반하는 강건한 과적합 현상을 조사한다. 본 연구는 이에 대응하기 위해 난이도 비례 레이블 스무딩 (DPLS)이라는 적응형 정규화 기법을 도입한다. DPLS는 측정된 각 학습 예제의 난이도에 정비례하는 레이블 스무딩 강도를 적용하여 암기 현상을 완화한다. 이러한 개입은 모델이 노이즈가 낀 특수성을 암기하는 대신, 더 보편적으로 적용 가능한 강건한 특징을 학습하도록 유도한다. 마지막으로, 5장에서는 방어적 학습에서 건설적 교정으로 관점을 전환하여, 대규모 비전-언어 모델의 고질적인 문제인 객체 환각 문제를 다룬다. 본 논문은 별도의 학습이 필요 없는 새로운 추론 방법론인 적대적 시각 대조 디코딩 (AVCD)을 제안한다. AVCD는 적대적 공격을 기발하게 재해석하여, 모델을 속이는 대신 환각을 최대로 유도하는 강력한 네거티브 시각 입력을 생성하는 데 사용한다. 이렇게 생성된 환각 편향 상태의 로짓 분포와 원본 이미지의 로짓 분포를 대조함으로써, AVCD는 시각적 근거가 없는 토큰을 정확하게 억제하고 추론 시 모델의 신뢰성을 크게 향상시킨다.

    결론적으로, 본 학위논문은 적대적 원리가 가진 놀라운 다재다능성을 입증한다. 먼저 데이터 인지 및 적응형 기법(DAFA, DPLS)으로 적대적 학습의 중대한 약점들을 보완하고, 나아가 이러한 원리를 추론 시점의 교정(AVCD)에 적용함으로써, 본 연구는 강력하고 다각적인 툴킷을 제공한다. 본 논문의 기여는 더 공정하고, 일반화 성능이 높으며, 궁극적으로 더 신뢰할 수 있는 AI 시스템을 구축하기 위한 실질적인 경로를 제시한다.
    번역하기

    인공지능(AI) 시스템이 학술적 호기심의 대상을 넘어 사회의 핵심 구성 요소로 자리 잡으면서, 자율주행이나 의료 진단과 같은 안전이 필수적인 도메인으로의 배치가 가속화되고 있다. 이러...

    인공지능(AI) 시스템이 학술적 호기심의 대상을 넘어 사회의 핵심 구성 요소로 자리 잡으면서, 자율주행이나 의료 진단과 같은 안전이 필수적인 도메인으로의 배치가 가속화되고 있다. 이러한 고위험 환경에서는 AI의 오류가 인간의 안녕과 사회적 신뢰에 심각한, 그리고 실질적인 영향을 미칠 수 있으므로, 모델의 신뢰성과 안정성을 확보하는 것은 타협 불가능한 최우선 목표가 된다. 이러한 노력의 핵심 기반 중 하나는 적대적 강건성으로, 이는 모델이 입력에 대해 작고 악의적으로 조작된 교란이 가해지는 상황에서도 정확하고 안정적인 예측을 유지하는 능력을 의미한다. 이 특성은 신뢰할 수 있는 AI를 구축하는 데 필수적이지만, 가장 보편적이고 기초적인 방어 기법인 적대적 학습은 완벽한 해결책과는 거리가 멀다. 본 학위논문은 표준적인 적대적 학습이 공정성 및 강건한 일반화 성능을 저해하는 등, 예기치 못한 미묘한 방식으로 모델의 안전성을 훼손할 수 있는 중대한 잠재적 결함들을 내포하고 있음을 주장한다. 본 연구는 이러한 난제들을 정면으로 다룬다. 이 중대한 결함들을 먼저 깊이 있게 조사하고, 나아가 적대적 공격의 핵심 원리를 모델의 교정 및 강화를 위한 건설적인 도구로 재활용하는 새로운 관점으로의 전환을 제시하며 이 분야의 발전에 기여하고자 한다.

    본 연구는 세 부분의 접근 방식을 제시한다. 첫째, 3장에서는 표준적인 적대적 학습이 특정 클래스의 강건성을 불균형적으로 저해하여 위험한 사각지대를 발생시키는 심각한 강건한 공정성 문제를 다룬다. 본 논문은 클래스 간의 유사성을 학습 과정에 민감하게 반영하는 새로운 방법론인 거리 인지 기반 공정 적대적 학습 (DAFA)을 제안한다. DAFA는 특성 공간 내 클래스 간의 기하학적 관계를 기반으로 클래스별 손실 가중치와 적응형 적대적 마진을 동적으로 할당한다. 이 메커니즘은 평균 성능의 저하 없이 취약 클래스의 강건성을 효과적으로 향상시켜, 보다 공평한 방어 모델을 구현한다.
    둘째, 4장에서는 모델이 적대적 학습 중 학습하기 어려운 예제들을 암기하기 시작하며, 이로 인해 역설적으로 보이지 않는 데이터에 대한 강건한 일반화 성능이 저해되는 직관에 반하는 강건한 과적합 현상을 조사한다. 본 연구는 이에 대응하기 위해 난이도 비례 레이블 스무딩 (DPLS)이라는 적응형 정규화 기법을 도입한다. DPLS는 측정된 각 학습 예제의 난이도에 정비례하는 레이블 스무딩 강도를 적용하여 암기 현상을 완화한다. 이러한 개입은 모델이 노이즈가 낀 특수성을 암기하는 대신, 더 보편적으로 적용 가능한 강건한 특징을 학습하도록 유도한다. 마지막으로, 5장에서는 방어적 학습에서 건설적 교정으로 관점을 전환하여, 대규모 비전-언어 모델의 고질적인 문제인 객체 환각 문제를 다룬다. 본 논문은 별도의 학습이 필요 없는 새로운 추론 방법론인 적대적 시각 대조 디코딩 (AVCD)을 제안한다. AVCD는 적대적 공격을 기발하게 재해석하여, 모델을 속이는 대신 환각을 최대로 유도하는 강력한 네거티브 시각 입력을 생성하는 데 사용한다. 이렇게 생성된 환각 편향 상태의 로짓 분포와 원본 이미지의 로짓 분포를 대조함으로써, AVCD는 시각적 근거가 없는 토큰을 정확하게 억제하고 추론 시 모델의 신뢰성을 크게 향상시킨다.

    결론적으로, 본 학위논문은 적대적 원리가 가진 놀라운 다재다능성을 입증한다. 먼저 데이터 인지 및 적응형 기법(DAFA, DPLS)으로 적대적 학습의 중대한 약점들을 보완하고, 나아가 이러한 원리를 추론 시점의 교정(AVCD)에 적용함으로써, 본 연구는 강력하고 다각적인 툴킷을 제공한다. 본 논문의 기여는 더 공정하고, 일반화 성능이 높으며, 궁극적으로 더 신뢰할 수 있는 AI 시스템을 구축하기 위한 실질적인 경로를 제시한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents iii
    • List of Tables vii
    • List of Figures viii
    • 1 INTRODUCTION 1
    • Abstract i
    • Contents iii
    • List of Tables vii
    • List of Figures viii
    • 1 INTRODUCTION 1
    • 2 BACKGROUND 7
    • 2.1 The Deep Learning Revolution in Computer Vision 7
    • 2.2 The Foundations of Adversarial Learning 8
    • 2.2.1 The Phenomenon of Adversarial Examples 9
    • 2.2.2 The Nature of Adversarial Attacks 10
    • 2.2.3 The Principles of Adversarial Defense 11
    • 2.3 Limitations of Standard Adversarial Training 12
    • 2.3.1 The Problem of Robust Fairness 12
    • 2.3.2 The Dynamics of Memorization and Hard Examples 12
    • 2.4 The Frontier of Multimodal AI: LVLMs and Their Reliability 13
    • 2.4.1 The Architecture of Large Vision-Language Models 13
    • 2.4.2 The Phenomenon of Vision Hallucination 14
    • 2.4.3 Inference-Time Interventions: The Rise of Contrastive Decoding 15
    • 2.5 Recent Advances and Future Directions 16
    • 2.5.1 Advances in Adversarial Machine Learning 16
    • 2.5.2 Evolving Learning Paradigms 17
    • 2.5.3 The Growing Importance of AI Reliability and Fairness 18
    • 2.6 Future Directions 19
    • 3 DAFA: Distance-Aware Fair Adversarial Training 21
    • 3.1 Introduction 21
    • 3.2 Preliminary 24
    • 3.3 Analysis 25
    • 3.3.1 Theoretical analysis 26
    • 3.3.2 Empirical verification 28
    • 3.4 Method 30
    • 3.4.1 Consideration of class-wise distance 31
    • 3.4.2 Distance-Aware fair adversarial training 33
    • 3.5 Experiment 36
    • 3.6 Conclusion 38
    • 3.7 Algorithms and Proofs 39
    • 3.7.1 The Algorithms of DAFA 39
    • 3.7.2 Proofs 41
    • 3.7.3 Theoretical analysis on adversarial margin 47
    • 3.8 Additional Experimental Results 49
    • 4 Regularizing Hard Examples Improves Adversarial Robustness 54
    • 4.1 Introduction 54
    • 4.2 Related work 57
    • 4.3 Analysis 58
    • 4.3.1 Analysis on the effect of hard examples 58
    • 4.3.2 Theoretical analysis on the training of hard examples 65
    • 4.3.3 Empirical verification on the training of hard examples 68
    • 4.4 Method 74
    • 4.4.1 Mitigation of the negative effect of hard examples 75
    • 4.4.2 Our method to mitigate the negative effect of hard examples 78
    • 4.5 Experiment 84
    • 4.6 Conclusion 87
    • 4.7 Proofs 88
    • 4.7.1 Proofs for Section 4.3.1 88
    • 4.7.2 Proofs for Section 4.3.2 90
    • 4.7.3 Proofs for Section 4.3.3 94
    • 4.7.4 Proofs for Section 4.4.2 95
    • 5 Visual Contrastive Decoding for Mitigating Hallucinations in Adversarial Large Vision Language Models 97
    • 5.1 Introduction 97
    • 5.2 Related Work 99
    • 5.3 Method 100
    • 5.3.1 Preliminaries 100
    • 5.3.2 Adversarial Visual Contrastive Decoding 102
    • 5.4 Analysis 104
    • 5.4.1 Hallucination Amplification Effect of Adversarial Perturbation 105
    • 5.4.2 Theoretical Analysis on Adversarial Perturbation 106
    • 5.4.3 Empirical Verification at the Encoder Level 107
    • 5.4.4 Empirical Verification at the LLM Level 108
    • 5.5 Experiment 110
    • 5.5.1 Experiment Settings 110
    • 5.5.2 Experiment Results 112
    • 5.5.3 Ablation Study on the Adversarial Loss 114
    • 5.5.4 Applicability to Black-Box Models via Adversarial Transfer 115
    • 5.6 Conclusion 117
    • 5.7 Algorithms and Proofs 117
    • 5.7.1 The Algorithms of AVCD 117
    • 5.7.2 Proof of the Theorem 119
    • 5.8 Additional Experimental Results 123
    • 6 CONCLUSION 135
    • 6.1 Summary 135
    • 6.2 Limitations 136
    • 6.3 Future Research 137
    • Bibliography 138
    • Abstract (In Korean) 151
    • 감사의 글 153
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼