RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    손실함수 이동법을 통한 강건한 마진기반 분류법 연구

    한글로보기

    https://www.riss.kr/link?id=T17550482

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    빅데이터 시대를 맞이하여 데이터의 수집 및 저장 비용은 낮아졌으나, 종속변수인 클래스 라벨에 오류가 포함되는 저품질 데이터(label noise) 문제는 여전히 이진 분류 모델의 신뢰성을 심각하게 위협하고 있다. 서포트 벡터 머신(SVM)이나 로지스틱 회귀와 같은 전통적인 마진 기반 분류 모형들은 수치해석적 최적화의 편의성을 담보하기 위해 볼록(convex)하고 유계가 없는(unbounded) 손실함수를 채택한다. 그러나 이러한 비유계성(unboundedness)으로 인해 라벨이 오염된 데이터는 손실함숫값을 크게 증가시켜, 결과적으로 올바른 분류함수의 탐색을 방해하고 결정 경계를 이상치(outlier) 방향으로 왜곡시키는 근본적인 원인이 된다. 기존의 절사(truncated) 손실함수나 비볼록(non-convex) 손실함수 기반의 접근법들은 최적화가 까다롭다는 한계가 존재한다. 본 연구에서는 이러한 문제를 해결하기 위해, 개별 관측치의 마진 크기에 따라 손실함수가 평행이동 되는 손실 이동 접근법(loss shifting approach)과 이를 규제하는 벌점화 프레임워크를 새롭게 제안한다. 훈련 자료의 각 관측치에 대하여 손실함수를 이동시키는 매개변수와 그 크기를 제어하는 맞춤형 벌점함수(penalty function) 를 도입함으로써, 기존의 비유계 손실함수를 유계된 형태를 가지는 유효 손실함수 (effective loss function)로 바꾼 후 최적화를 통해 모수를 추정한다. 특히, 본 연구는 마진을 위한 증폭 조절(Amplified Thresholding for Margin; ATM) 규칙을 고안하여 메커니즘을 명확히 수식화하였다. 이는 기존에 제안되었던 다양한 로버스트 분류 기법들을 하나의 공통된 체계 내에서 포괄할 수 있는 통합적 프레임워크를 제공한다. 제안된 통합 프레임워크의 효율적인 추정을 위해, 본 연구는 분류함수와 이동 매개변수를 순차적으로 업데이트하는 교대 적합 알고리즘을 설계하였다. 제안된 방법론은 다양한 라벨 오염비율 환경에서도 기존의 고전적인 분류 모델들보다 뛰어난 예측 성능을 입증하였다.
    번역하기

    빅데이터 시대를 맞이하여 데이터의 수집 및 저장 비용은 낮아졌으나, 종속변수인 클래스 라벨에 오류가 포함되는 저품질 데이터(label noise) 문제는 여전히 이진 분류 모델의 신뢰성을 심...

    빅데이터 시대를 맞이하여 데이터의 수집 및 저장 비용은 낮아졌으나, 종속변수인 클래스 라벨에 오류가 포함되는 저품질 데이터(label noise) 문제는 여전히 이진 분류 모델의 신뢰성을 심각하게 위협하고 있다. 서포트 벡터 머신(SVM)이나 로지스틱 회귀와 같은 전통적인 마진 기반 분류 모형들은 수치해석적 최적화의 편의성을 담보하기 위해 볼록(convex)하고 유계가 없는(unbounded) 손실함수를 채택한다. 그러나 이러한 비유계성(unboundedness)으로 인해 라벨이 오염된 데이터는 손실함숫값을 크게 증가시켜, 결과적으로 올바른 분류함수의 탐색을 방해하고 결정 경계를 이상치(outlier) 방향으로 왜곡시키는 근본적인 원인이 된다. 기존의 절사(truncated) 손실함수나 비볼록(non-convex) 손실함수 기반의 접근법들은 최적화가 까다롭다는 한계가 존재한다. 본 연구에서는 이러한 문제를 해결하기 위해, 개별 관측치의 마진 크기에 따라 손실함수가 평행이동 되는 손실 이동 접근법(loss shifting approach)과 이를 규제하는 벌점화 프레임워크를 새롭게 제안한다. 훈련 자료의 각 관측치에 대하여 손실함수를 이동시키는 매개변수와 그 크기를 제어하는 맞춤형 벌점함수(penalty function) 를 도입함으로써, 기존의 비유계 손실함수를 유계된 형태를 가지는 유효 손실함수 (effective loss function)로 바꾼 후 최적화를 통해 모수를 추정한다. 특히, 본 연구는 마진을 위한 증폭 조절(Amplified Thresholding for Margin; ATM) 규칙을 고안하여 메커니즘을 명확히 수식화하였다. 이는 기존에 제안되었던 다양한 로버스트 분류 기법들을 하나의 공통된 체계 내에서 포괄할 수 있는 통합적 프레임워크를 제공한다. 제안된 통합 프레임워크의 효율적인 추정을 위해, 본 연구는 분류함수와 이동 매개변수를 순차적으로 업데이트하는 교대 적합 알고리즘을 설계하였다. 제안된 방법론은 다양한 라벨 오염비율 환경에서도 기존의 고전적인 분류 모델들보다 뛰어난 예측 성능을 입증하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In the era of big data, the cost of data collection and storage has substantially decreased. However, the problem of low-quality data, particularly label noise in class labels serving as the response variable, continues to seriously threaten the reliability of binary classification models. Traditional margin-based classification models, such as support vector machines (SVMs) and logistic regression, employ convex and unbounded loss functions to ensure the computational convenience of numerical optimization. However, due to this unboundedness, label-contaminated observations can induce excessively large loss values, thereby obstructing the search for an appropriate classification function and fundamentally distorting the decision boundary toward outliers. Existing approaches based on truncated loss functions or non-convex loss functions have limitations in that their optimization procedures are computationally challenging.
    In this study, to address these issues, we propose a loss shifting approach in which the loss function is horizontally shifted according to the margin size of each individual observation, together with a penalized framework that regulates the amount
    of shifting. By introducing an observation-specific shifting parameter for the loss function and a customized penalty function that controls its magnitude, the proposed method transforms a conventional unbounded loss function into an effective loss function with a bounded form, and estimates the model parameters through optimization. In particular, this study develops an Amplified Thresholding for Margin (ATM)rule, thereby providing a clear mathematical formulation of the proposed mechanism. This offers a unified framework that can encompass various existing robust classification methods within a common structure.
    For efficient estimation under the proposed unified framework, we design an alternating fitting algorithm that sequentially updates the classification function and the shifting parameters. The proposed methodology demonstrates superior predictive performance compared with conventional classification models under various levels of label contamination.
    번역하기

    In the era of big data, the cost of data collection and storage has substantially decreased. However, the problem of low-quality data, particularly label noise in class labels serving as the response variable, continues to seriously threaten the relia...

    In the era of big data, the cost of data collection and storage has substantially decreased. However, the problem of low-quality data, particularly label noise in class labels serving as the response variable, continues to seriously threaten the reliability of binary classification models. Traditional margin-based classification models, such as support vector machines (SVMs) and logistic regression, employ convex and unbounded loss functions to ensure the computational convenience of numerical optimization. However, due to this unboundedness, label-contaminated observations can induce excessively large loss values, thereby obstructing the search for an appropriate classification function and fundamentally distorting the decision boundary toward outliers. Existing approaches based on truncated loss functions or non-convex loss functions have limitations in that their optimization procedures are computationally challenging.
    In this study, to address these issues, we propose a loss shifting approach in which the loss function is horizontally shifted according to the margin size of each individual observation, together with a penalized framework that regulates the amount
    of shifting. By introducing an observation-specific shifting parameter for the loss function and a customized penalty function that controls its magnitude, the proposed method transforms a conventional unbounded loss function into an effective loss function with a bounded form, and estimates the model parameters through optimization. In particular, this study develops an Amplified Thresholding for Margin (ATM)rule, thereby providing a clear mathematical formulation of the proposed mechanism. This offers a unified framework that can encompass various existing robust classification methods within a common structure.
    For efficient estimation under the proposed unified framework, we design an alternating fitting algorithm that sequentially updates the classification function and the shifting parameters. The proposed methodology demonstrates superior predictive performance compared with conventional classification models under various levels of label contamination.

    더보기

    목차 (Table of Contents)

    • 1 서론 1
    • 2 이동 모수를 이용한 손실 최소화 6
    • 2.1 이동 모수의 성질 7
    • 2.2 강건 회귀와의 연관성 13
    • 1 서론 1
    • 2 이동 모수를 이용한 손실 최소화 6
    • 2.1 이동 모수의 성질 7
    • 2.2 강건 회귀와의 연관성 13
    • 3 이동 모수, 벌점 함수, 그리고 유효 손실 16
    • 4 학습 알고리즘 30
    • 5 마진 조절 기법 33
    • 5.1 마진을 위한 증폭 조절 (Amplified Thresholding for Margin; ATM) 33
    • 6 실험 설계 및 결과 분석 38
    • 6.1 실험의 목적 38
    • 6.2 실험 요인과 비교 구조 38
    • 6.3 손실 이동 모형의 학습 절차 39
    • 6.4 모의실험 설계 41
    • 6.4.1 선형 결정 경계 41
    • 6.4.2 비선형 부호 패리티 구조 42
    • 6.5 실제 자료 분석 설계 43
    • 6.6 모의실험 결과 45
    • 6.6.1 선형 모의실험: ntrain = 100 45
    • 6.6.2 선형 모의실험: ntrain = 200 47
    • 6.6.3 비선형 모의실험 50
    • 6.7 실제 자료 분석 결과 53
    • 6.8 종합적 해석 58
    • 참고문헌 59
    • 부록 A. 벌점함수와 유효 손실의 형태 62
    • 부록 B. 분류함수 f 업데이트를 위한 부문제 70
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼