금융 신용 데이터는 보통 심각한 클래스 불균형, long tailed·outlier 분포, 중복·분산된 이벤트 집계 구조, 그리고 규제 환경에서의 높은 설명 가능성 요구라는 복합적인 난제를 동시에 가진다. ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
금융 신용 데이터는 보통 심각한 클래스 불균형, long tailed·outlier 분포, 중복·분산된 이벤트 집계 구조, 그리고 규제 환경에서의 높은 설명 가능성 요구라는 복합적인 난제를 동시에 가진다. ...
금융 신용 데이터는 보통 심각한 클래스 불균형, long tailed·outlier 분포, 중복·분산된 이벤트 집계 구조, 그리고 규제 환경에서의 높은 설명 가능성 요구라는 복합적인 난제를 동시에 가진다. 일반적으로 사용되는 min–max 스케일링, 트리 기반 앙상블, 단일 테이블 기반 딥러닝 모델과 같은 표준 접근법은 이러한 특성을 모두 반영하지 못해, 정보 손실과 표현력 한계, 계산 비용 문제를 야기한다. 본 연구에서는 Gauss Rank Transformation 전처리, 상관·의미 기반 패칭(correlation/semantic patching)을 적용한 Tabular Transformer, prototype voting head로 구성된 이단계 학습 프레임워크를 제안한다. 제안한 구조는 (i) long-tailed·outlier 분포를 갖는 연속형 변수의 스케일 왜곡을 완화하고, (ii) 중복·분산된 이벤트 집계를 정보 손실 없이 패치 단위 표현으로 통합하며, (iii) 변수 간 관계를 self-attention을 통해 직접적으로 모델링하고, (iv) 프로토타입 수준의 투표 구조를 통해 개별 예측에 대한 직관적인 설명을 제공한다. PFCT의 대규모 실제 신용 데이터에 대한 실험에서, 제안 모델은 Gradient Boosting 및 최신 tabular 딥러닝 모델 대비 KS, 상위 p% 구간 recall, AUC 등 모든 지표에서 일관된 성능 향상을 보였으며, 연체 확률의 수준과 중요 특성 조합을 동시에 해석할 수 있는 실용적인 설명 가능성을 확인하였다.
다국어 초록 (Multilingual Abstract)
Financial credit data typically exhibit severe class imbalance, long-tailed and outlier-prone distributions, duplicated and fragmented event aggregation structures, and strong explainability requirements under regulatory constraints. Standard approach...
Financial credit data typically exhibit severe class imbalance, long-tailed and outlier-prone distributions, duplicated and fragmented event aggregation structures, and strong explainability requirements under regulatory constraints. Standard approaches such as min–max scaling, tree-based ensembles, and single-table deep learning models fail to fully accommodate these characteristics, leading to information loss, limited representational capacity, and increased computational cost. In this study, we propose a two-stage learning framework that combines preprocessing with Gauss Rank Transformation, a Tabular Transformer with correlation/semantic patching, and a
prototype voting head. The proposed architecture (i) mitigates scale distortion in continuous variables with long - tailed, outlier-rich distributions, (ii) integrates duplicated and fragmented event aggregates into patch-level representations without information loss, (iii) explicitly models relationships among variables via self-attention, and (iv) provides intuitive, case-level explanations through a prototype-level voting structure. In experiments on a large-scale real-world credit dataset from PFCT, the proposed model consistently outperforms Gradient Boosting and recent state-of-the-art tabular deep learning models across all metrics, including KS, recall in the top-p% risk segment, and AUC, while offering practical interpretability that jointly captures both the level of default probability and the key feature combinations driving each prediction.
목차 (Table of Contents)