
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
특징 래핑을 통한 숫자형 특징과 범주형 특징이 혼합된 데이터의 클래스 분류 성능 향상 기법
이재성(Jaesung Lee),김대원(Daewon Kim) 한국정보과학회 2009 정보과학회논문지 : 소프트웨어 및 응용 Vol.36 No.12
본 논문에서는 혼합형 데이터에 대한 특징 선별 기법의 효율성을 비교하기 위해 특징 필터링과 특징 래핑을 통한 특징 선별 후, 클래스 분류 성능을 측정하였다. 혼합형 데이터는 숫자형 특징과 범주형 특징이 함께 혼합되어 있으므로, 숫자형 특징을 범주형 특징으로 이산화를 하여 단일형 데이터로 변환한 뒤 특징 선별 기법 등을 적용할 수 있다. 본 연구에서는 혼합형 데이터를 전처리하여 단일형 데이터로 변환하고, 널리 활용되는 특징 필터링 기법과 특징 래핑 기법을 통해 클래스 분류 성능을 높일 수 있는 특징 집합을 선별하였다. 선별된 특징 집합을 통한 클래스 분류 성능을 비교한 결과, 특징 필터링에 비해 특징 래핑을 통해 선별한 특징 집합을 활용하여 클래스 분류를 하였을 때 분류 정확도가 높은 것을 확인할 수 있었다. In this letter, we evaluate the classification performance of mixed numeric and categorical data for comparing the efficiency of feature filtering and feature wrapping. Because the mixed data is composed of numeric and categorical features, the feature selection method was applied to data set after discretizing the numeric features in the given data set. In this study, we choose the feature subset for improving the classification performance of the data set after preprocessing. The experimental result of comparing the classification performance show that the feature wrapping method is more reliable than feature filtering method in the aspect of classification accuracy.
음성 기반 치매 조기 진단 연구의 확장 : Mel-Spectrogram 중심 접근의 한계와 언어적 특징 기반 다중모달 분석
장관종 국제차세대융합기술학회 2026 차세대융합기술학회논문지 Vol.10 No.3
고령화 사회의 가속화로 치매 조기 선별의 중요성이 강조되는 가운데, 음성 기반 인공지능은 비침습적·저비용 대안으로 주목받고 있다. 선행 연구에서는 음성 신호를 멜 스펙트로그램(Mel-Spectrogram)으로 변환하여 CNN, ViT 모델 등에 적용하였으나, 분류 정확도가 약 61~62% 수준에 머물며 음향적 특징만으로는 치매 특유의 인지 저하를 포착하는 데 구조적 한계가 있음을 확인하였다. 본 연구는 이러한 한계를 극복하기 위해 ADRESS-2020 데이터셋을 기반으로 전사 텍스트에서 추출한 언어적 특징을 결합한 다중모달 분석 접근을 제안하였다. 연구 결과, 어휘다양성, 문장 복잡도, 의미적 응집성 등 14개의 언어적 변수만으로도 교차검증 정확도 76.8%를 달성하며 선행 연구의 음향 기반 모델 성능을 크게 상회하였다. 특히 특징 수 대비 성능 효율성 측면에서 언어적 특징은 고차원 딥러닝특징보다 월등히 높은 수치를 기록하여, 소규모 의료 데이터 환경에서 특징의 질적 설계가 중요함을 입증하였다. 음향과 언어 특징의 결합은 안정성과 분산 측면에서 가장 균형 잡힌 결과를 나타냈으나, 모든 특징을 결합한 고차원환경에서는 차원의 저주로 인한 성능 저하가 관찰되었다. 모델 비교에서는 로지스틱 회귀가 가장 우수한 일반화 성능을 보였으며, 이는 실제 임상 현장에서 해석 가능하고 단순한 모델의 실용성이 높음을 시사한다. 본 연구는 언어적특징 중심의 다중모달 분석이 치매 조기 선별의 정확성과 신뢰성을 높이는 핵심 전략임을 실증하였다. As the acceleration of population aging intensifies the importance of early dementia screening, voice-based artificial intelligence is gaining significant attention as a non-invasive and cost-effective alternative. A previous study utilized voice signals converted into Mel-spectrograms and applied them to models such as CNN and ViT, but found that classification accuracy remained at approximately 61–62%, confirming structural limitations in capturing dementia-specific cognitive decline using only acoustic features . To overcome these limitations, this study proposes a multimodal analysis approach that integrates linguistic features extracted from transcribed text using the ADRESS-2020 dataset. The experimental results demonstrated that just 14 linguistic variables—including lexical diversity, syntactic complexity, and semantic coherence—achieved a cross-validation accuracy of 76.8%, significantly outperforming the acoustic-based models from the previous study. In terms of performance efficiency relative to the number of features, linguistic features recorded substantially higher values than high-dimensional deep learning features, proving that qualitative feature engineering is crucial in small-scale medical data environments. While the combination of acoustic and linguistic features yielded the most balanced results in terms of stability and variance, a performance decline due to the "curse of dimensionality" was observed when all high-dimensional features were combined. In model comparisons, logistic regression exhibited the most superior generalization performance, suggesting that simple, interpretable models are more practical for real-world clinical settings. This study empirically validates that multimodal analysis centered on linguistic features is a core strategy for enhancing the accuracy and reliability of early dementia screening.
모바일 결제 시스템의 수요 예측을 위한 신경망에서 특징 선별 기법
김호준,조윤석,김경미 한국정보과학회 2018 정보과학회논문지 Vol.45 No.4
In this paper, we present a time series prediction technique based on neural network as a methodology for forecasting service demand of mobile payment system. We propose a two-stage neural network model for the feature selection process and the prediction process. Three types of fuzzy membership functions were adopted for the representation of feature data, and a hyperbox-based neural network model is used for the evaluation of feature relevance factor. The proposed feature selection technique reduces the amount of computation and eliminates erroneous feature data in the learning data set. We evaluated the usefulness of the proposed method through experiments using two years of data obtained form actual smart campus systems. 본 논문에서는 모바일 결제시스템의 서비스 수요예측을 위한 방법론으로서 신경망 기반의 시계열예측 기법을 제시한다. 예측에 필요한 특징 선별과정과 시계열 데이터의 예측과정을 위하여 2단계 신경망 모델을 제안하며 그 동작 특성과 알고리즘에 관해 기술한다. 특징 데이터의 표현을 위하여 3종류의 퍼지 멤버쉽함수를 적용하며, 하이퍼박스 기반의 신경망 모델을 사용하여 특징의 연관도 요소를 평가하는 방법을 제시한다. 제안된 특징 선별 기법은 예측 시스템의 계산량을 감소시키며, 학습데이터 집합에서 왜곡된 특징 데이터를 제거할 수 있게 한다. 실제 스마트캠퍼스 시스템에서 취득한 2년간의 데이터를 사용하여 실험을 수행하고 그 결과를 통하여 제안된 기법의 유용성을 평가한다.
유전자 알고리즘과 Feature Wrapping을 통한 마이크로어레이 데이타 중복 특징 소거법
이재성(Jae-sung Lee),김대원(Dae-won Kim) 한국정보과학회 2008 정보과학회논문지 : 소프트웨어 및 응용 Vol.35 No.8
Due to the high dimensional problem, typically machine learning algorithms have relied on feature selection techniques in order to perform effective classification in microarray gene expression datasets. However, the large number of features compared to the number of samples makes the task of feature selection computationally inprohibitive and prone to errors. One of traditional feature selection approach was feature filtering; measuring one gene per one step. Then feature filtering was an univariate approach that cannot validate multivariate correlations. In this paper, we proposed a function for measuring both class separability and correlations. With this approach, we solved the problem related to feature filtering approach. 본 논문에서는 유전자 사이의 상관계수가 높은 마이크로어레이 데이타에 대하여 제안하는 알고리즘을 통해 상관계수가 낮은 유전자들의 부집합을 만들고, 이에 대해 적합 함수를 통한 평가로 기존 방법론이 가지는 한계를 극복할 수 있도록 하였다. 기존 방법론은 개별 특징의 평가를 통해 중복 특징을 제거하며, 상관계수에 대한 고려가 없어 선택된 유전자 부집합들의 상관계수가 높은 문제가 있었다. 이에 따라 제안하는 알고리즘은 특징간의 관계를 평가하는 Feature Wrapping 기법을 활용하여, 추출된 유전자 부집합에 포함된 유전자 사이의 상관관계가 낮고, 클래스 구분력이 높은 특징을 갖도록 하였다.
다중 레이블 인문 데이터의 효과적인 특징 선별을 위한 이주 개체 정제연산 기반 다중 개체군 유전알고리즘
박민우 ( Park Minwoo ),이재성 ( Lee Jaesung ) 중앙대학교 인공지능인문학연구소 2021 인공지능인문학연구 Vol.7 No.-
Multi-label feature selection is a preprocessing method that can be used to analyze, for example multi-label humanity data. In particular, a multi-population genetic algorithm is verified to exhibit a better performance for identifying an appropriate subset compared with existing genetic algorithms in that a variety of populations was preserved, and premature convergence was prevented. However, with this method, the inflow of closely related features to multi-labels is unlikely to search for the solution. This study proposes an effective multi-population genetic algorithm for multi-label feature selection. In the proposed method, a multi-population genetic algorithm with a refinement process in migrated individuals maintains a variety of populations, promotes the inflow of features closely related to multi-labels, and ultimately enhances the search performance. Experimental results indicate that the proposed method exhibit better performance than the compared multi-population algorithms.
다중 레이블 데이터 분류를 위한 상호 정보 척도를 이용한 특징 선별 기법
임현기,김대원 한국정보과학회 2012 정보과학회논문지 : 소프트웨어 및 응용 Vol.39 No.10
Lately multi-label data set occurs in many applications. However it is difficult to apply in machine learning and data mining fields. There are two reasons: One is that most of researches are focusing on the single-label problem and the other is that the previous methods do not account the characteristics of multi-label. Existing methods cannot be applied to multi-label data because most of feature selection methods have focused in single-label data. For applying existing method, there have been used label transformation methods. However label transformation may lead to information loss of data. In this paper, we propose feature selection method for multi-label data considering the dependency between labels. We experimented classification for demonstrating the superiority of proposed method. This shows that the proposed method is better than previous feature selection methods. 최근 많은 응용에서 다중 레이블 데이터가 발생하고 있다. 하지만 이 데이터는 기존 기계 학습, 데이터 마이닝 분야의 방법 적용이 어렵다. 그 이유는 크게 두 가지로 기존 방법들이 단일 레이블 데이터에 초점을 맞추고 있다는 것과 다중 레이블 데이터의 특성을 반영하지 못하고 있다는 것이다. 대부분의 특징 선별 기법은 단일 레이블 데이터에 초점을 맞추고 있기 때문에 다중 레이블 데이터에는 기존 특징 선별 기법들을 적용할 수 없다. 다중 레이블 데이터에 특징 선별 기법을 적용하기 위해서 다중 레이블 데이터를 단일 레이블 데이터로 전환하는 방법들이 사용된다. 하지만 레이블 변환은 데이터 고유의 특성을 반영하지 못하고 정보 손실을 가져올 수 있다. 본 논문은 레이블과 레이블 사이의 연관성을 고려하여 다중 레이블 데이터에 바로 적용할 수 있는 특징 선별 기법을 제안한다. 제안하는 방법의 우수성을 보이기 위해 클래스 분류 실험을 하였다. 이를 통해 기존 특징 선별 기법들에 비해서 제안하는 기법의 성능이 우수하다는 것을 보였다.
NSGA-II 알고리즘을 이용한 다중 레이블 분류 문제에서 특징 선별
윤정훈(Jeonghun Yoon),이재성(Jaesung Lee),김대원(Dae-Won Kim) 한국정보과학회 2013 정보과학회논문지 : 소프트웨어 및 응용 Vol.40 No.3
최근 하나 이상의 클래스 레이블을 가지는 데이터에 대한 클래스 분류 기법들이 연구되고 있다. 그 중 몇몇 연구들에서 다중 레이블 분류의 성능을 높이기 위해 특징 선별 기법을 사용했다. 그러나 다중 레이블 데이터의 복잡성으로 인해 기존의 특징 선별 기법을 적용하는 것은 성능 향상에 한계가 있다. 본 논문은 다목적 최적화 알고리즘인 NSGA-II를 다중 레이블 특징 선별 문제에 활용했다. 또한 본 논문에서는 특징 선별 문제에 적합한 유전자 조작 기법을 제안하여 NSGA-II에 적용했다. 그리고 실제 다중 레이블 데이터 셋에서 실험을 통해 제안하는 알고리즘이 기존 유전자 알고리즘을 사용한 특징 선별 기법보다 더 좋은 성능을 가지는 것을 보인다. Recently, a lot of researchers are interested in multi-label classification. Some of the researchers use feature subset selection to improve performance in multi-label classification. However, because multi-label problem is more complex than single-label, single-label feature selection algorithms have limitation to be applied to multi-label data. In this paper, NSGA-II, which is multi-objective optimize algorithm, is used for multi-label feature selection. In addition, this study proposes a novel genetic operator for feature selection problem. Finally, experiments on real world data set show that proposed algorithm achieves better performance than traditional genetic algorithm.
회귀 최적화 기반 관련도와 중복도를 이용한 다중 레이블 특징 선별 기법
임현기 한국컴퓨터정보학회 2024 한국컴퓨터정보학회논문지 Vol.29 No.11
High-dimensional data causes difficulties in machine learning due to high time consumption and large memory requirements. In particular, in a multi-label environment, higher complexity is required as much as the number of labels. This paper proposes a feature selection method to improve classification performance in multi-label settings. The method considers three types of relationships: between features, between features and labels, and between labels themselves. To achieve this, a regression-based objective function is designed. This objective function calculates the linear relationships between features and labels and uses mutual information to compute relationships between features and between labels. By minimizing this objective function, the optimal weights for feature selection are found. To optimize the objective function, a gradient descent method is applied to develop a fast-converging algorithm. The experimental results on six multi-label datasets show that the proposed method outperforms existing multi-label feature selection techniques. The classification performance of the proposed method, averaged over six datasets, showed a Hamming loss of 0.1285, a ranking loss of 0.1811, and a multi-label accuracy of 0.6416. Compared to the AMI(Approximating Mutual Information) algorithm, the performance was better by 0.0148, 0.0435, and 0.0852, respectively.