머신러닝에서 분류 모델은 분류 오류를 최소화하고 예측 정확도를 높이기 위한 모형이다. 그러나 실제 상황에서는 특정 클래스의 발생빈도가 매우 낮은 불균형 데이터가 많이 나타난다. 이...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T16974202
서울 : 성균관대학교 일반대학원, 2024
학위논문(석사) -- 성균관대학교 일반대학원 , 통계학과 , 2024. 2
2024
한국어
불균형 데이터 ; SVDD ; 베이지안 방법 ; minor class
서울
Bayesian support vector data description using minor class
43 p. : 삽화 ; 30 cm
지도교수: 김재직
참고문헌: p. 40-41
I804:11040-000000177391
0
상세조회0
다운로드머신러닝에서 분류 모델은 분류 오류를 최소화하고 예측 정확도를 높이기 위한 모형이다. 그러나 실제 상황에서는 특정 클래스의 발생빈도가 매우 낮은 불균형 데이터가 많이 나타난다. 이...
머신러닝에서 분류 모델은 분류 오류를 최소화하고 예측 정확도를 높이기 위한 모형이다. 그러나 실제 상황에서는 특정 클래스의 발생빈도가 매우 낮은 불균형 데이터가 많이 나타난다. 이러한 불균형 데이터는 효과적인 분류 모델을 학습시키기 어렵다는 문제점이 있다. 문제를 해결하기 위해 최근 많은 연구가 이루어지고 있으며, 이를 해결하기 위한 방법들은 데이터와 알고리즘 두 가지 측면의 접근법으로 크게 나눌 수 있다. 데이터 측면에서는 리샘플링을 통해 데이터 분포를 재조정하거나 다수 클래스의 관측값을 줄여서 문제를 해결한다. 알고리즘 측면에서는 불균형 데이터를 고려하여 분류 모델을 수정하거나 앙상블 기법을 활용하여 분류 결과를 결합한다. 이러한 방법들 중에서 SVDD (Support Vector Data Description)는 원-클래스 러닝 (one-class learning) 방법으로서 다수 클래스의 중심과 경계선을 구하는 방식으로 클래스를 식별한다. 그러나 SVDD를 포함한 대부분의 방법론은 모델이 갖는 불확실성을 고려하지 않아 신뢰하기 어렵다. 또는, SVDD에 있어 다수 클래스만 고려하여 다수 클래스의 중심과 경계선을 찾기 보다는 소수 클래스의 정보까지 이용한다면 불균형 데이터의 분류에 있어 더 정확성을 올릴 수 있다. 따라서 본 논문은 불균형 데이터를 분류하기 위해 소수 클래스의 데이터를 추가로 이용하고 불확실성을 고려한 베이지안 SVDD 방법을 제안한다. 제안하는 방법은 SVDD에서 소수 클래스의 데이터를 추가로 사용하고, 불확실성을 고려하기 위해 베이지안 접근법을 사용한다. 본 논문에서 제안하는 모형의 성능을 검증하기 위해 다양한 상황에서의 모의실험을 진행한다.
다국어 초록 (Multilingual Abstract)
In machine learning, a classification model aims to minimize classification errors and enhance prediction accuracy. However, there are often involve imbalanced data in real-world, where the occurrence frequency of a specific class is significantly low...
In machine learning, a classification model aims to minimize classification errors and enhance prediction accuracy. However, there are often involve imbalanced data in real-world, where the occurrence frequency of a specific class is significantly low. Dealing with such imbalanced data poses challenges in developing effective classification models. A lot of research has been conducted recently to solve the problem, and methods to solve this problem can be broadly divided into two approaches : data and algorithm. On the data approach, the issue can be addressed by adjusting data distribution through resampling or reducing observations in major classe. On the algorithmic approach, modify classification models to account for unbalanced data or use ensemble techniques to combine classification results. Among these methods, Support Vector Data Description (SVDD) is a one-class learning approach that focuses on determining the center of major class by calculating the boundary line. However, most methods, including SVDD, do not consider the uncertainty of the model and therefore difficult to trust. In SVDD, the center and boundary of the major class are determined without accounting for information from the minor class. Additional use of data from the minor class can enhance accuracy in classifying imbalanced data. Therefore, this thesis proposes a bayesian SVDD method that additionally uses minor class data and and considers uncertainty to classify imbalanced data. The proposed method additionally uses minor class data from SVDD and uses a bayesian approach to consider uncertainty. To verify the performance of the method proposed in this thesis, simulation experiments are conducted in various situations.
목차 (Table of Contents)