기계 학습에서 데이터 품질은 모든 측면의 근간이 되며, 오프라인 및 온라인 학습 환경 모두에서 모델 학습의 핵심 요소에 크게 영향을 미칩니다. 본 논문은 이러한 맥락에서 데이터 품질 문...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
기계 학습에서 데이터 품질은 모든 측면의 근간이 되며, 오프라인 및 온라인 학습 환경 모두에서 모델 학습의 핵심 요소에 크게 영향을 미칩니다. 본 논문은 이러한 맥락에서 데이터 품질 문...
기계 학습에서 데이터 품질은 모든 측면의 근간이 되며, 오프라인 및 온라인 학습 환경 모두에서 모델 학습의 핵심 요소에 크게 영향을 미칩니다. 본 논문은 이러한 맥락에서 데이터 품질 문제를 해결하여 학습 결과를 향상시키기 위한 전략을 탐구 하며, 이는 지능형 시스템 발전에 필수적입니다.
논문의 첫 번째 부분에서는 노이즈가 포함된 회귀 데이터 처리를 위한 오프라인 전략을 연구합니다. 레이블과 특징 간의 연속적이고 정렬된 상관 관계라는 본질적 속성을 활용하여, 유사한 레이블이 밀접하게 관련된 특징과 대응한다는 점에 착안 한 새로운 이웃 정렬 혼합 모델(neighborhood-aligning mixture model)을 제안하고, 이를 통해 노이즈 데이터의 영향을 효과적으로 식별하고 완화합니다.
두 번째 부분에서는 온라인 학습 환경에서의 노이즈 필터링으로 초점을 전환합 니다. 먼저, 노이즈가 망각을 악화시키는 방법을 온라인 지속 학습 환경에서 입증 하며, 데이터를 깨끗하게 유지하기 위해 중심성 기반 확률 그래프 앙상블을 활용한 새로운 솔루션인 Self-Purified Replay를 제안합니다. 추가적으로, 온라인 멀티모달 비디오 데이터의 필터링 문제를 다루며, 작업 관련성과 정보성을 기준으로 샘플을 필터링하는 방법을 제안합니다. 이러한 접근법은 전체 멀티모달 데이터를 저장할 필요를 제거하며, 실시간으로 노이즈 없는 맥락 적응을 가능하게 합니다.
마지막으로, 본 연구의 기여를 요약하고 데이터 품질을 개선하여 학습 성능을 지원하기 위한 향후 연구 방향을 논의하며 논문을 마무리합니다.
다국어 초록 (Multilingual Abstract)
Data quality plays a foundational role in all aspects of machine learning, significantly impacting core aspects of model learning in both offline and online training environments. This dissertation investigates strategies to enhance learning outcomes ...
Data quality plays a foundational role in all aspects of machine learning, significantly impacting core aspects of model learning in both offline and online training environments. This dissertation investigates strategies to enhance learning outcomes by addressing data quality challenges in these contexts, which are crucial for the advancement of intelligent systems.
In the first part of this work, we investigate offline strategies for handling noisy regression data. By leveraging the intrinsic property of continuous and ordered correlations between labels and features—where similar labels correspond to closely related features—we propose a novel neighborhood-aligning mixture model to effectively identify and mitigate the impact of noisy data.
The second part of this dissertation shifts focus to noise filtering in online learning environments. We first emphasize its importance in the online continual learning setting by demonstrating how noise exacerbates forgetting and propose a novel solution, Self-Purified Replay, which leverages centrality-based stochastic graph ensembles to maintain a clean subset of data. Additionally, we address the challenge of filtering online multimodal video data by filtering towards task-relevant and informative samples. This approach eliminates the need for full multimodal data storage while enabling real-time, noise-free, in-context adaptations.
Finally, we conclude by summarizing the contributions of this research and discussing promising directions for future work aimed at enhancing data quality to support improved learning.
목차 (Table of Contents)
Resources: Deep Learning with Mini Whiteboards
Teachers TV Teachers TVHow to Get A+ in Your Online Classes: Startegies fSuccessful Online Learning
한서대학교 Shaneil DipasupilDeep Learning with TensorFlow
K-MOOC 선문대학교 김종혁, 이성철딥러닝(Deep Learning)이란
신한대학교 신종우비전공자를 위한 AI 딥러닝(Deep Learning)
K-MOOC 한국과학기술원 오종훈