RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    트랜스포머 구조를 이용한 연체확률 예측 모델 = Applying Transformers for Delinquency Probabilities Prediction

    한글로보기

    https://www.riss.kr/link?id=T17451934

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    금융 신용 데이터는 보통 심각한 클래스 불균형, long tailed·outlier 분포, 중복·분산된 이벤트 집계 구조, 그리고 규제 환경에서의 높은 설명 가능성 요구라는 복합적인 난제를 동시에 가진다. 일반적으로 사용되는 min–max 스케일링, 트리 기반 앙상블, 단일 테이블 기반 딥러닝 모델과 같은 표준 접근법은 이러한 특성을 모두 반영하지 못해, 정보 손실과 표현력 한계, 계산 비용 문제를 야기한다. 본 연구에서는 Gauss Rank Transformation 전처리, 상관·의미 기반 패칭(correlation/semantic patching)을 적용한 Tabular Transformer, prototype voting head로 구성된 이단계 학습 프레임워크를 제안한다. 제안한 구조는 (i) long-tailed·outlier 분포를 갖는 연속형 변수의 스케일 왜곡을 완화하고, (ii) 중복·분산된 이벤트 집계를 정보 손실 없이 패치 단위 표현으로 통합하며, (iii) 변수 간 관계를 self-attention을 통해 직접적으로 모델링하고, (iv) 프로토타입 수준의 투표 구조를 통해 개별 예측에 대한 직관적인 설명을 제공한다. PFCT의 대규모 실제 신용 데이터에 대한 실험에서, 제안 모델은 Gradient Boosting 및 최신 tabular 딥러닝 모델 대비 KS, 상위 p% 구간 recall, AUC 등 모든 지표에서 일관된 성능 향상을 보였으며, 연체 확률의 수준과 중요 특성 조합을 동시에 해석할 수 있는 실용적인 설명 가능성을 확인하였다.
    번역하기

    금융 신용 데이터는 보통 심각한 클래스 불균형, long tailed·outlier 분포, 중복·분산된 이벤트 집계 구조, 그리고 규제 환경에서의 높은 설명 가능성 요구라는 복합적인 난제를 동시에 가진다. ...

    금융 신용 데이터는 보통 심각한 클래스 불균형, long tailed·outlier 분포, 중복·분산된 이벤트 집계 구조, 그리고 규제 환경에서의 높은 설명 가능성 요구라는 복합적인 난제를 동시에 가진다. 일반적으로 사용되는 min–max 스케일링, 트리 기반 앙상블, 단일 테이블 기반 딥러닝 모델과 같은 표준 접근법은 이러한 특성을 모두 반영하지 못해, 정보 손실과 표현력 한계, 계산 비용 문제를 야기한다. 본 연구에서는 Gauss Rank Transformation 전처리, 상관·의미 기반 패칭(correlation/semantic patching)을 적용한 Tabular Transformer, prototype voting head로 구성된 이단계 학습 프레임워크를 제안한다. 제안한 구조는 (i) long-tailed·outlier 분포를 갖는 연속형 변수의 스케일 왜곡을 완화하고, (ii) 중복·분산된 이벤트 집계를 정보 손실 없이 패치 단위 표현으로 통합하며, (iii) 변수 간 관계를 self-attention을 통해 직접적으로 모델링하고, (iv) 프로토타입 수준의 투표 구조를 통해 개별 예측에 대한 직관적인 설명을 제공한다. PFCT의 대규모 실제 신용 데이터에 대한 실험에서, 제안 모델은 Gradient Boosting 및 최신 tabular 딥러닝 모델 대비 KS, 상위 p% 구간 recall, AUC 등 모든 지표에서 일관된 성능 향상을 보였으며, 연체 확률의 수준과 중요 특성 조합을 동시에 해석할 수 있는 실용적인 설명 가능성을 확인하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Financial credit data typically exhibit severe class imbalance, long-tailed and outlier-prone distributions, duplicated and fragmented event aggregation structures, and strong explainability requirements under regulatory constraints. Standard approaches such as min–max scaling, tree-based ensembles, and single-table deep learning models fail to fully accommodate these characteristics, leading to information loss, limited representational capacity, and increased computational cost. In this study, we propose a two-stage learning framework that combines preprocessing with Gauss Rank Transformation, a Tabular Transformer with correlation/semantic patching, and a
    prototype voting head. The proposed architecture (i) mitigates scale distortion in continuous variables with long - tailed, outlier-rich distributions, (ii) integrates duplicated and fragmented event aggregates into patch-level representations without information loss, (iii) explicitly models relationships among variables via self-attention, and (iv) provides intuitive, case-level explanations through a prototype-level voting structure. In experiments on a large-scale real-world credit dataset from PFCT, the proposed model consistently outperforms Gradient Boosting and recent state-of-the-art tabular deep learning models across all metrics, including KS, recall in the top-p% risk segment, and AUC, while offering practical interpretability that jointly captures both the level of default probability and the key feature combinations driving each prediction.
    번역하기

    Financial credit data typically exhibit severe class imbalance, long-tailed and outlier-prone distributions, duplicated and fragmented event aggregation structures, and strong explainability requirements under regulatory constraints. Standard approach...

    Financial credit data typically exhibit severe class imbalance, long-tailed and outlier-prone distributions, duplicated and fragmented event aggregation structures, and strong explainability requirements under regulatory constraints. Standard approaches such as min–max scaling, tree-based ensembles, and single-table deep learning models fail to fully accommodate these characteristics, leading to information loss, limited representational capacity, and increased computational cost. In this study, we propose a two-stage learning framework that combines preprocessing with Gauss Rank Transformation, a Tabular Transformer with correlation/semantic patching, and a
    prototype voting head. The proposed architecture (i) mitigates scale distortion in continuous variables with long - tailed, outlier-rich distributions, (ii) integrates duplicated and fragmented event aggregates into patch-level representations without information loss, (iii) explicitly models relationships among variables via self-attention, and (iv) provides intuitive, case-level explanations through a prototype-level voting structure. In experiments on a large-scale real-world credit dataset from PFCT, the proposed model consistently outperforms Gradient Boosting and recent state-of-the-art tabular deep learning models across all metrics, including KS, recall in the top-p% risk segment, and AUC, while offering practical interpretability that jointly captures both the level of default probability and the key feature combinations driving each prediction.

    더보기

    목차 (Table of Contents)

    • 제 1 장 서 론 1
    • 제 2 장 이론적 배경 5
    • 제 1 절 정형 데이터와 트리 기반 모델 (GBDT) 5
    • 제 2 절 딥러닝과 트랜스포머 (Transformer) 9
    • 제 3 절 어텐션 매커니즘의 원리 11
    • 제 1 장 서 론 1
    • 제 2 장 이론적 배경 5
    • 제 1 절 정형 데이터와 트리 기반 모델 (GBDT) 5
    • 제 2 절 딥러닝과 트랜스포머 (Transformer) 9
    • 제 3 절 어텐션 매커니즘의 원리 11
    • 제 4 절 정형 데이터를 위한 트랜스포머 연구 13
    • 제 3 장 모 델 15
    • 제 1 절 데이터 전처리 16
    • 제 2 절 Tabular Transformer 19
    • 제 3 절 Voting Head 30
    • 제 4 장 실험 결과 36
    • 제 1 절 평가 지표 36
    • 제 2 절 실험 결과 37
    • 제 5 장 결론 및 향후 연구 과제 39
    • 참고문헌 41
    • 상세 실험 결과 43
    • 표 목차
    • [표 1] 전처리 방법에 따른 성능 비교 19
    • [표 2] Patching 방법에 따른 성능 비교 25
    • [표 3] Tabular Transformer의 inference 방법에 따른 성능 비교 29
    • [표 4] Voting Head 도입에 따른 성능 비교 35
    • [표 5] Voting Head를 추가한 뒤 end-to-end 학습 방법에 따른 성능 비교 36
    • [표 6] 타 모델들과의 성능 비교 37
    • 그림 목차
    • [그림 1] Self Attention의 구조도 12
    • [그림 2] 모델의 전체 개요 15
    • [그림 3] Min-Max Scaling과 Gauss Rank Transformation의 비교 18
    • [그림 4] Correlation Patching 생성 과정 24
    • [그림 5] 입력 임베딩 구조 26
    • [그림 6] Prototype Voting Head의 동작 메커니즘 34
    • [그림 7] KS그래프와 AUC그래프 36
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼