RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Minimal-Feedback Learning for Generative AI = 생성형 인공지능을 위한 최소 피드백 학습

    한글로보기

    https://www.riss.kr/link?id=T17450623

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Modern generative AI systems, including diffusion models and large language models (LLMs), have achieved remarkable capabilities but require extensive human supervision for task-specific alignment. This thesis presents a novel paradigm of minimal-feedback learning that dramatically reduces the human annotation burden while maintaining alignment quality. We make two primary contributions. First, we demonstrate that only 3 minutes of binary human feedback suffices to censor unwanted visual concepts from pre-trained diffusion models. Using reward model ensembles and guided sampling, we reduce the generation of malign images from up to 68% to below 1% across four diverse censoring tasks without model retraining. Second, we extend the minimal-feedback principle to text generation through adversarial self-play, where LLMs iteratively refine their own guard prompts by alternating between attacker and defender roles. We provide a unified theoretical framework that casts both adversarial refinement and iterative self-feedback methods (including PromptWizard and ProTeGi) as fixed-point iterations, establishing convergence guarantees for black-box prompt optimization. Our cross-modal results demonstrate that treating human time—not compute—as the scarcest resource enables practical, accessible AI alignment. This work opens new directions for efficient model steering across modalities, making responsible AI deployment feasible for resource-constrained researchers and organizations.
    번역하기

    Modern generative AI systems, including diffusion models and large language models (LLMs), have achieved remarkable capabilities but require extensive human supervision for task-specific alignment. This thesis presents a novel paradigm of minimal-feed...

    Modern generative AI systems, including diffusion models and large language models (LLMs), have achieved remarkable capabilities but require extensive human supervision for task-specific alignment. This thesis presents a novel paradigm of minimal-feedback learning that dramatically reduces the human annotation burden while maintaining alignment quality. We make two primary contributions. First, we demonstrate that only 3 minutes of binary human feedback suffices to censor unwanted visual concepts from pre-trained diffusion models. Using reward model ensembles and guided sampling, we reduce the generation of malign images from up to 68% to below 1% across four diverse censoring tasks without model retraining. Second, we extend the minimal-feedback principle to text generation through adversarial self-play, where LLMs iteratively refine their own guard prompts by alternating between attacker and defender roles. We provide a unified theoretical framework that casts both adversarial refinement and iterative self-feedback methods (including PromptWizard and ProTeGi) as fixed-point iterations, establishing convergence guarantees for black-box prompt optimization. Our cross-modal results demonstrate that treating human time—not compute—as the scarcest resource enables practical, accessible AI alignment. This work opens new directions for efficient model steering across modalities, making responsible AI deployment feasible for resource-constrained researchers and organizations.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    확산 모델과 대규모 언어 모델(LLM)을 포함한 현대 생성형 AI 시스템은 놀라운 성능을 달성했지만, 작업별 정렬을 위해서는 방대한 인간 감독이 필요하다. 본 논문은 정렬 품질을 유지하면서도 인간 주석 부담을 획기적으로 줄이는 최소 피드백 학습(minimal-feedback learning)이라는 새로운 패러다임을 제시한다. 본 연구의 주요 기여는 다음과 같다. 첫째, 사전 학습된 확산 모델에서 원치 않는 시각적 개념을 검열하는 데 단 3분의 이진 인간 피드백만으로 충분함을 입증한다. 보상 모델 앙상블과 유도 샘플링을 사용하여, 모델 재훈련 없이 네 가지 다양한 검열 작업에서 악성 이미지 생성률을 최대 68%에서 1% 미만으로 감소시켰다. 둘째, 최소 피드백 원리를 텍스트 생성으로 확장하여, LLM이 공격자와 방어자 역할을 번갈아 수행하며 자체 보호 프롬프트를 반복적으로 개선하는 적대적 자기 학습(adversarial self-play) 방법을 제안한다. 우리는 적대적 개선과 반복적 자기 피드백 방법(PromptWizard와 ProTeGi 포함)을 고정점 반복으로 해석하는 통합 이론 프레임워크를 제공하며, 블랙박스 프롬프트 최적화에 대한 수렴 보장을 확립한다. 본 연구의 교차 모달 결과는 컴퓨팅 자원이 아닌 인간의 시간을 가장 희소한 자원으로 취급함으로써 실용적이고 접근 가능한 AI 정렬이 가능함을 보여준다. 이 연구는 모달리티 전반에 걸친 효율적인 모델 조향의 새로운 방향을 열어, 자원이 제한된 연구자와 조직들도 안정적인 AI 배포를 가능하게 한다.
    번역하기

    확산 모델과 대규모 언어 모델(LLM)을 포함한 현대 생성형 AI 시스템은 놀라운 성능을 달성했지만, 작업별 정렬을 위해서는 방대한 인간 감독이 필요하다. 본 논문은 정렬 품질을 유지하면서...

    확산 모델과 대규모 언어 모델(LLM)을 포함한 현대 생성형 AI 시스템은 놀라운 성능을 달성했지만, 작업별 정렬을 위해서는 방대한 인간 감독이 필요하다. 본 논문은 정렬 품질을 유지하면서도 인간 주석 부담을 획기적으로 줄이는 최소 피드백 학습(minimal-feedback learning)이라는 새로운 패러다임을 제시한다. 본 연구의 주요 기여는 다음과 같다. 첫째, 사전 학습된 확산 모델에서 원치 않는 시각적 개념을 검열하는 데 단 3분의 이진 인간 피드백만으로 충분함을 입증한다. 보상 모델 앙상블과 유도 샘플링을 사용하여, 모델 재훈련 없이 네 가지 다양한 검열 작업에서 악성 이미지 생성률을 최대 68%에서 1% 미만으로 감소시켰다. 둘째, 최소 피드백 원리를 텍스트 생성으로 확장하여, LLM이 공격자와 방어자 역할을 번갈아 수행하며 자체 보호 프롬프트를 반복적으로 개선하는 적대적 자기 학습(adversarial self-play) 방법을 제안한다. 우리는 적대적 개선과 반복적 자기 피드백 방법(PromptWizard와 ProTeGi 포함)을 고정점 반복으로 해석하는 통합 이론 프레임워크를 제공하며, 블랙박스 프롬프트 최적화에 대한 수렴 보장을 확립한다. 본 연구의 교차 모달 결과는 컴퓨팅 자원이 아닌 인간의 시간을 가장 희소한 자원으로 취급함으로써 실용적이고 접근 가능한 AI 정렬이 가능함을 보여준다. 이 연구는 모달리티 전반에 걸친 효율적인 모델 조향의 새로운 방향을 열어, 자원이 제한된 연구자와 조직들도 안정적인 AI 배포를 가능하게 한다.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Scope and Contributions 2
    • 1.3 Thesis Structure
    • 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Scope and Contributions 2
    • 1.3 Thesis Structure
    • 2 Censored Sampling of Diffusion Models Using 3 Minutes of Human Feedback (Reprinted from NeurIPS 2023) 5
    • 2.1 Introduction 5
    • 2.2 Background on diffusion probabilistic models 9
    • 2.3 Problem description: Censored sampling with human feedback 10
    • 2.4 Reward model and human feedback 12
    • 2.5 Reward model ensemble for benign-dominant setups 12
    • 2.6 Imitation learning for malign-dominant setups 14
    • 2.7 Transfer learning for time-independent reward 15
    • 2.8 Sampling 15
    • 2.9 Experiments 17
    • 2.10 MNIST: Censoring 7 with a strike-through cross 17
    • 2.11 LSUN church: Censoring watermarks from latent diffusion model 18
    • 2.12 ImageNet: Tench (fish) without human faces 19
    • 2.13 LSUN bedroom: Censoring broken bedrooms 21
    • 2.14 Stable Diffusion: Censoring unexpected embedded texts 22
    • 2.15 Conclusion 24
    • 3 Robust Prompt Refinement via Adversarial Self-Play and Iterative Self-Feedback Framework 25
    • 3.1 Introduction 25
    • 3.2 Adversarial Self-Play for Robust Prompt Refinement 30
    • 3.2.1 Problem Setting and Threat Model 30
    • 3.2.2 Adversarial Self-Play Algorithm 32
    • 3.2.3 Simulation via LLM-Based Agents 35
    • 3.3 Iterative Self-Feedback Prompt Optimization: A Unified Framework 37
    • 3.3.1 General Loop Operators and Feedback Mechanisms 37
    • 3.3.2 Convergence, Fixed Points, and Theoretical Guarantees 41
    • 4 Conclusion 46
    • 4.1 Summary of Findings 46
    • 4.2 Limitations 46
    • 4.3 Future Work 47
    • 4.4 Closing Remark 47
    • A Broader impacts & safety 48
    • B Limitations 49
    • C Human subject and evaluation 50
    • D Prior Works 51
    • E GUI interface 54
    • F Reward model: Further details 56
    • G Backward guidance and recurrence 58
    • G.1 Backward guidance 58
    • G.2 Recurrence 59
    • H MNIST crossed 7: Experiment details and image samples 60
    • H.1 Diffusion model 60
    • H.2 Reward model and training 61
    • H.3 Sampling and ablation study 61
    • H.4 Censored generation samples 62
    • I LSUN church: Experiment details and image samples 66
    • I.1 Pre-trained diffusion model 66
    • I.2 Malign image definition 66
    • I.3 Reward model training 67
    • I.4 Sampling and ablation study 68
    • I.5 Censored generation samples 68
    • J ImageNet tench: Experiment details and image samples 72
    • J.1 Pre-trained diffusion model 72
    • J.2 Reward model training 72
    • J.3 Sampling and ablation study 73
    • J.4 Censored generation samples 73
    • K LSUN bedroom: Experiment details and image samples 77
    • K.1 Pre-trained diffusion model 77
    • K.2 Malign image definition 77
    • K.3 Reward model training 78
    • K.4 Sampling and ablation study 80
    • K.5 Censored generation samples 81
    • L Stable Diffusion: Experiment details and image samples 95
    • L.1 Pre-trained diffusion model 95
    • L.2 Reward model training 95
    • L.3 Sampling and ablation study 96
    • L.4 Censored generation samples 96
    • M Transfer learning ablation 100
    • N Using malign images from secondary source 101
    • Bibliography 104
    • 초록 112
    • Acknowledgements 113
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼