RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Towards Model-based Data Selection for Noisy and Irrelevant Data = 노이즈와 불필요한 데이터에 대한 모델 기반 데이터 선택 방법

    한글로보기

    https://www.riss.kr/link?id=T17315320

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    기계 학습에서 데이터 품질은 모든 측면의 근간이 되며, 오프라인 및 온라인 학습 환경 모두에서 모델 학습의 핵심 요소에 크게 영향을 미칩니다. 본 논문은 이러한 맥락에서 데이터 품질 문제를 해결하여 학습 결과를 향상시키기 위한 전략을 탐구 하며, 이는 지능형 시스템 발전에 필수적입니다.

    논문의 첫 번째 부분에서는 노이즈가 포함된 회귀 데이터 처리를 위한 오프라인 전략을 연구합니다. 레이블과 특징 간의 연속적이고 정렬된 상관 관계라는 본질적 속성을 활용하여, 유사한 레이블이 밀접하게 관련된 특징과 대응한다는 점에 착안 한 새로운 이웃 정렬 혼합 모델(neighborhood-aligning mixture model)을 제안하고, 이를 통해 노이즈 데이터의 영향을 효과적으로 식별하고 완화합니다.

    두 번째 부분에서는 온라인 학습 환경에서의 노이즈 필터링으로 초점을 전환합 니다. 먼저, 노이즈가 망각을 악화시키는 방법을 온라인 지속 학습 환경에서 입증 하며, 데이터를 깨끗하게 유지하기 위해 중심성 기반 확률 그래프 앙상블을 활용한 새로운 솔루션인 Self-Purified Replay를 제안합니다. 추가적으로, 온라인 멀티모달 비디오 데이터의 필터링 문제를 다루며, 작업 관련성과 정보성을 기준으로 샘플을 필터링하는 방법을 제안합니다. 이러한 접근법은 전체 멀티모달 데이터를 저장할 필요를 제거하며, 실시간으로 노이즈 없는 맥락 적응을 가능하게 합니다.

    마지막으로, 본 연구의 기여를 요약하고 데이터 품질을 개선하여 학습 성능을 지원하기 위한 향후 연구 방향을 논의하며 논문을 마무리합니다.
    번역하기

    기계 학습에서 데이터 품질은 모든 측면의 근간이 되며, 오프라인 및 온라인 학습 환경 모두에서 모델 학습의 핵심 요소에 크게 영향을 미칩니다. 본 논문은 이러한 맥락에서 데이터 품질 문...

    기계 학습에서 데이터 품질은 모든 측면의 근간이 되며, 오프라인 및 온라인 학습 환경 모두에서 모델 학습의 핵심 요소에 크게 영향을 미칩니다. 본 논문은 이러한 맥락에서 데이터 품질 문제를 해결하여 학습 결과를 향상시키기 위한 전략을 탐구 하며, 이는 지능형 시스템 발전에 필수적입니다.

    논문의 첫 번째 부분에서는 노이즈가 포함된 회귀 데이터 처리를 위한 오프라인 전략을 연구합니다. 레이블과 특징 간의 연속적이고 정렬된 상관 관계라는 본질적 속성을 활용하여, 유사한 레이블이 밀접하게 관련된 특징과 대응한다는 점에 착안 한 새로운 이웃 정렬 혼합 모델(neighborhood-aligning mixture model)을 제안하고, 이를 통해 노이즈 데이터의 영향을 효과적으로 식별하고 완화합니다.

    두 번째 부분에서는 온라인 학습 환경에서의 노이즈 필터링으로 초점을 전환합 니다. 먼저, 노이즈가 망각을 악화시키는 방법을 온라인 지속 학습 환경에서 입증 하며, 데이터를 깨끗하게 유지하기 위해 중심성 기반 확률 그래프 앙상블을 활용한 새로운 솔루션인 Self-Purified Replay를 제안합니다. 추가적으로, 온라인 멀티모달 비디오 데이터의 필터링 문제를 다루며, 작업 관련성과 정보성을 기준으로 샘플을 필터링하는 방법을 제안합니다. 이러한 접근법은 전체 멀티모달 데이터를 저장할 필요를 제거하며, 실시간으로 노이즈 없는 맥락 적응을 가능하게 합니다.

    마지막으로, 본 연구의 기여를 요약하고 데이터 품질을 개선하여 학습 성능을 지원하기 위한 향후 연구 방향을 논의하며 논문을 마무리합니다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Data quality plays a foundational role in all aspects of machine learning, significantly impacting core aspects of model learning in both offline and online training environments. This dissertation investigates strategies to enhance learning outcomes by addressing data quality challenges in these contexts, which are crucial for the advancement of intelligent systems.

    In the first part of this work, we investigate offline strategies for handling noisy regression data. By leveraging the intrinsic property of continuous and ordered correlations between labels and features—where similar labels correspond to closely related features—we propose a novel neighborhood-aligning mixture model to effectively identify and mitigate the impact of noisy data.

    The second part of this dissertation shifts focus to noise filtering in online learning environments. We first emphasize its importance in the online continual learning setting by demonstrating how noise exacerbates forgetting and propose a novel solution, Self-Purified Replay, which leverages centrality-based stochastic graph ensembles to maintain a clean subset of data. Additionally, we address the challenge of filtering online multimodal video data by filtering towards task-relevant and informative samples. This approach eliminates the need for full multimodal data storage while enabling real-time, noise-free, in-context adaptations.

    Finally, we conclude by summarizing the contributions of this research and discussing promising directions for future work aimed at enhancing data quality to support improved learning.
    번역하기

    Data quality plays a foundational role in all aspects of machine learning, significantly impacting core aspects of model learning in both offline and online training environments. This dissertation investigates strategies to enhance learning outcomes ...

    Data quality plays a foundational role in all aspects of machine learning, significantly impacting core aspects of model learning in both offline and online training environments. This dissertation investigates strategies to enhance learning outcomes by addressing data quality challenges in these contexts, which are crucial for the advancement of intelligent systems.

    In the first part of this work, we investigate offline strategies for handling noisy regression data. By leveraging the intrinsic property of continuous and ordered correlations between labels and features—where similar labels correspond to closely related features—we propose a novel neighborhood-aligning mixture model to effectively identify and mitigate the impact of noisy data.

    The second part of this dissertation shifts focus to noise filtering in online learning environments. We first emphasize its importance in the online continual learning setting by demonstrating how noise exacerbates forgetting and propose a novel solution, Self-Purified Replay, which leverages centrality-based stochastic graph ensembles to maintain a clean subset of data. Additionally, we address the challenge of filtering online multimodal video data by filtering towards task-relevant and informative samples. This approach eliminates the need for full multimodal data storage while enabling real-time, noise-free, in-context adaptations.

    Finally, we conclude by summarizing the contributions of this research and discussing promising directions for future work aimed at enhancing data quality to support improved learning.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Acknowledgements xxiii
    • Chapter 1 Introduction 1
    • 1.1 Thesis Overview 2
    • Abstract i
    • Acknowledgements xxiii
    • Chapter 1 Introduction 1
    • 1.1 Thesis Overview 2
    • Chapter 2 Background 5
    • 2.1 Preliminaries 5
    • 2.1.1 Self-supervised Learning 5
    • 2.1.2 Online Learning 6
    • 2.1.3 Continual Learning 7
    • 2.1.4 Learning with Noisy Labels 8
    • 2.1.5 Learning from Regression Data 10
    • 2.2 Related Works 11
    • 2.2.1 Related Works in Noisy Labels 11
    • 2.2.2 Related Works for Multimodal Data Selection 14
    • Chapter 3 Continual Learning on Noisy Data Streams via Self-Purified Replay 18
    • 3.1 Introduction 18
    • 3.2 Problem Statement 21
    • 3.2.1 Noisy Labeled Continual Learning 21
    • 3.2.2 Motivation: Noise induced Amnesia 21
    • 3.3 Approach to Noisy Labeled Continual Learning 22
    • 3.3.1 Self-Replay 23
    • 3.3.2 Self-Centered Filter 25
    • 3.4 Experiments 32
    • 3.4.1 Experimental Design 32
    • 3.4.2 Baselines 34
    • 3.4.3 Training Details 35
    • 3.4.4 Results and Analysis 37
    • 3.5 Summary 47
    • Chapter 4 Sample Selection via Contrastive Fragmentation for Noisy Label Regression 53
    • 4.1 Introduction 53
    • 4.2 ConFrag: Contrastive Fragmentation 56
    • 4.2.1 Contrastive Fragment Pairing 57
    • 4.2.2 Training Feature Extractors for Contrastive Pairs 59
    • 4.2.3 Mixture of Neighboring Fragments 60
    • 4.2.4 Neighborhood Jittering 63
    • 4.3 Theory of ConFrag 63
    • 4.3.1 Classification versus Regression for Feature Learning 64
    • 4.3.2 Fragmentation and Neighborhood Jittering 64
    • 4.3.3 Prediction Depth Analysis 66
    • 4.4 Experiments 68
    • 4.4.1 Settings 69
    • 4.4.2 Results and Discussion 79
    • 4.5 Summary 98
    • Chapter 5 ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams 100
    • 5.1 Introduction 100
    • 5.2 Method 103
    • 5.2.1 Multimodal Alignment Filtering 104
    • 5.2.2 Relevance Filtering 105
    • 5.2.3 Specificity Filtering 107
    • 5.3 Experiments 109
    • 5.3.1 Settings 109
    • 5.3.2 Baselines 111
    • 5.3.3 Training Details 113
    • 5.3.4 Results 113
    • 5.3.5 Analysis & Ablations 114
    • 5.4 Summary 124
    • Chapter 6 Conclusion 130
    • 6.1 Summary 130
    • 6.2 Future Work 131
    • 요약 162
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼