RISS 학술연구정보서비스

검색
다국어 입력

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

변환된 중국어를 복사하여 사용하시면 됩니다.

예시)
  • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
  • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
닫기
    인기검색어 순위 펼치기

    RISS 인기검색어

      검색결과 좁혀 보기

      선택해제
      • 좁혀본 항목 보기순서

        • 원문유무
        • 원문제공처
        • 등재정보
        • 학술지명
          펼치기
        • 주제분류
        • 발행연도
          펼치기
        • 작성언어

      오늘 본 자료

      • 오늘 본 자료가 없습니다.
      더보기
      • 무료
      • 기관 내 무료
      • 유료
      • KCI등재

        다중 레이블 데이터 분류를 위한 상호 정보 척도를 이용한 특징 선별 기법

        임현기,김대원 한국정보과학회 2012 정보과학회논문지 : 소프트웨어 및 응용 Vol.39 No.10

        Lately multi-label data set occurs in many applications. However it is difficult to apply in machine learning and data mining fields. There are two reasons: One is that most of researches are focusing on the single-label problem and the other is that the previous methods do not account the characteristics of multi-label. Existing methods cannot be applied to multi-label data because most of feature selection methods have focused in single-label data. For applying existing method, there have been used label transformation methods. However label transformation may lead to information loss of data. In this paper, we propose feature selection method for multi-label data considering the dependency between labels. We experimented classification for demonstrating the superiority of proposed method. This shows that the proposed method is better than previous feature selection methods. 최근 많은 응용에서 다중 레이블 데이터가 발생하고 있다. 하지만 이 데이터는 기존 기계 학습, 데이터 마이닝 분야의 방법 적용이 어렵다. 그 이유는 크게 두 가지로 기존 방법들이 단일 레이블 데이터에 초점을 맞추고 있다는 것과 다중 레이블 데이터의 특성을 반영하지 못하고 있다는 것이다. 대부분의 특징 선별 기법은 단일 레이블 데이터에 초점을 맞추고 있기 때문에 다중 레이블 데이터에는 기존 특징 선별 기법들을 적용할 수 없다. 다중 레이블 데이터에 특징 선별 기법을 적용하기 위해서 다중 레이블 데이터를 단일 레이블 데이터로 전환하는 방법들이 사용된다. 하지만 레이블 변환은 데이터 고유의 특성을 반영하지 못하고 정보 손실을 가져올 수 있다. 본 논문은 레이블과 레이블 사이의 연관성을 고려하여 다중 레이블 데이터에 바로 적용할 수 있는 특징 선별 기법을 제안한다. 제안하는 방법의 우수성을 보이기 위해 클래스 분류 실험을 하였다. 이를 통해 기존 특징 선별 기법들에 비해서 제안하는 기법의 성능이 우수하다는 것을 보였다.

      • KCI등재

        효과적인 애스팩트 마이닝을 위한 다중 레이블 분류접근법

        원종윤,이건창 한국경영정보학회 2020 Information systems review Vol.22 No.3

        최근의 감성분류 연구는 출력변수가 하나인 단일레이블 분류방법을 사용한 연구가 많다. 특히, 이러한 연구는 하나의 극성 값(긍정, 부정)만을 찾는 연구가 많다. 그러나 한 문장 안에는 다중적인 의미가 내포되어 있다. 그 중에서도 감정과 오피니언이 이러한 특징을 갖는다. 본 논문은 두 가지 연구목적을 제시한다. 첫째, 한 문장 안에 다양한 토픽(주제 또는 애스팩트)이 있다는 사실을 기반으로, 해당 문장을 각 애스팩트 별로 감성을 분류하는 애스팩트 마이닝을 수행한다. 둘째, 두개 이상의 종속변수(출력 값)를 한 번에 분석하는 다중레이블 분류방법을 적용한다. 이에 본 연구는 감성분류의 연구가 단일분류기에 의해서만 이루어진 연구를 개선하고자 다중레이블 분류방법에 의한 애스팩트 마이닝을 수행하고자 한다. 이와 같은 연구목적을 달성하기 위해 국내 뮤지컬 데이터를 수집하였다. 분석결과 문장 안에 있는 다양한 애스팩트별 감성을 추출하였고, 유의한 결과를 얻었다.

      • KCI등재

        CNN 기반 HTTPS 트래픽의 서비스 및 응용 단위 다중 레이블 분류

        김보선,최정우,박지태,백의준,김명섭 한국통신학회 2023 韓國通信學會論文誌 Vol.48 No.5

        웹 응용 프로그램은 HTTP 프로토콜의 보안 문제로 HTTPS 프로토콜을 사용한다. HTTPS 응용 트래픽 분류를위해 다양한 연구가 진행되고 있으나, 암호화로 인해 분류하기 어렵다. 이를 해결하기 위한 방법들은 multi-class 분류이며 서비스 단위로 응용 트래픽을 분류한다. 웹 서비스는 여러 응용 프로그램의 조합으로 구성되며, 트래픽또한 여러 응용 프로그램의 조합으로 하나의 서비스를 이룬다. 기존 연구와 같이 서비스 단위로만 분류할 경우, 서비스 내 여러 응용 프로그램의 트래픽이 섞여 미탐 혹은 오탐할 수 있다. 따라서 우리는 multi-label 분류를 사용하여 플로우 별로 서비스 단위인 주라벨과 응용 프로그램 단위인 부라벨을 정의하고, 합성 곱 신경망을 설계하여 6개의 서비스와 서비스 내 14개의 응용 프로그램을 분류한다. 6개의 서비스 분류 정확도는 100%, 14개의 응용 프로그램 분류 정확도는 85% 성능으로 서비스 단위와 응용 프로그램 단위로 트래픽을 분류하였다.

      • KCI등재

        NSGA-II 알고리즘을 이용한 다중 레이블 분류 문제에서 특징 선별

        윤정훈(Jeonghun Yoon),이재성(Jaesung Lee),김대원(Dae-Won Kim) 한국정보과학회 2013 정보과학회논문지 : 소프트웨어 및 응용 Vol.40 No.3

        최근 하나 이상의 클래스 레이블을 가지는 데이터에 대한 클래스 분류 기법들이 연구되고 있다. 그 중 몇몇 연구들에서 다중 레이블 분류의 성능을 높이기 위해 특징 선별 기법을 사용했다. 그러나 다중 레이블 데이터의 복잡성으로 인해 기존의 특징 선별 기법을 적용하는 것은 성능 향상에 한계가 있다. 본 논문은 다목적 최적화 알고리즘인 NSGA-II를 다중 레이블 특징 선별 문제에 활용했다. 또한 본 논문에서는 특징 선별 문제에 적합한 유전자 조작 기법을 제안하여 NSGA-II에 적용했다. 그리고 실제 다중 레이블 데이터 셋에서 실험을 통해 제안하는 알고리즘이 기존 유전자 알고리즘을 사용한 특징 선별 기법보다 더 좋은 성능을 가지는 것을 보인다. Recently, a lot of researchers are interested in multi-label classification. Some of the researchers use feature subset selection to improve performance in multi-label classification. However, because multi-label problem is more complex than single-label, single-label feature selection algorithms have limitation to be applied to multi-label data. In this paper, NSGA-II, which is multi-objective optimize algorithm, is used for multi-label feature selection. In addition, this study proposes a novel genetic operator for feature selection problem. Finally, experiments on real world data set show that proposed algorithm achieves better performance than traditional genetic algorithm.

      • KCI등재

        머신러닝 기반의 기업 리뷰 다중 분류 : 부분 문법 적용을 중심으로

        백혜연,장영균 한국경영정보학회 2023 경영정보학연구 Vol.25 No.3

        최근 많은 분야에서 기계학습에 대한 연구가 활발히 진행되고 있는데, 상당수의 연구들이 학습 모델의 성능을 개선하는 최신 방법론을 제시하고 있다. 본 연구에서는 방법론의 개발 못지않게 기계학습에 투입되는 훈련용 데이터의 ‘품질’을 개선하는 것 역시 중요하다는 점에 착안하여, 코퍼스 분석에서 자주 사용되는 ‘부분 문법’ 처리 프로세스를 통해 훈련 데이터의 품질을 향상시키는 방법을 제시한다. 우리나라 100대 기업에 근무하는 재직자들이 채용플랫폼에 게시하는 방대한 양의 비정형 기업 리뷰 텍스트 데이터를 수집하고, 데이터 품질을 부분 문법 프로세스로 개선한 후, 부분 문법이 적용된 분류 모델이 적용되지 않은 모델보다 분류 성능이 우수함을 확인하였다. 분류 카테고리는 직원 몰입의 5가지 요인으로 상정하였는데, 국내 직장인들이 기업 리뷰가 각 유형별로 빈도에 차이가 있는지를 분석하였다. 추가로 리뷰 양상이 코로나 팬데믹 전후로 어떠한 변화가 있었는지도 분석하였다. 본 연구를 통해 국내 직장인들의 생생한 일터 경험들을 자동적으로 식별하고 분류하여, 이직을 포함한 주요한 조직문화 현상의 행태와 유발 원인 등을 유추해 볼 수 있는 근거를 제공한다. Unlike the previous works focusing on the state-of-the-art methodologies to improve the performance of machine learning models, this study improves the ’quality' of training data used in machine learning. We propose a method to enhance the quality of training data through the processing of 'local grammar,' frequently used in corpus analysis. We collected a vast amount of unstructured corporate review text data posted by employees working in the top 100 companies in Korea. After improving the data quality using the local grammar process, we confirmed that the classification model with local grammar outperformed the model without it in terms of classification performance. We defined five factors of work engagement as classification categories, and analyzed how the pattern of reviews changed before and after the COVID-19 pandemic. Through this study, we provide evidence that shows the value of the local grammar-based automatic identification and classification of employee experiences, and offer some clues for significant organizational cultural phenomena.

      • KCI등재후보

        국가별 행정체계 특성을 반영한 인공지능 활용 해외 주소데이터 품질검증 기법

        김진실,이경희,조완섭 사)한국빅데이터학회 2022 한국빅데이터학회 학회지 Vol.7 No.2

        글로벌 시대에 들어서면서 수입식품 안전관리에 대한 중요성이 증가하고 있다. 해외 식품업체 주소정보는 수입식품 안전관리를 위한 핵심 정보로써 식품위해 발생시 신속한 대처와 사후관리를 위해 반드시 검증되어야 한다. 그러나 각국의 주소체계가 다른 관계로 하나의 검증시스템이 모든 국가의 주소를 검증할 수는 없다. 또한, 주소검증은 사용하는 분야에 따라 검정목적이 상이할 수 있다. 본 논문에서는 주어진 해외 식품업체 주소로부터 해당 국가의 행정구역 레벨로 분류하는 문제를 다룬다. 수입식품 안전관리를 정확하고 효율적으로 하기 위하여 수입식품제조업체 주소를 해당 국가의 행정구역 수준으로 정확하게 매칭하는 것이 필요하다. 수입식품이 생산⋅제조되는 위치와 식품제조에 영향을 줄 수 있는 환경정보, 재난재해 정보를 결합함으로써 선제적 수입식품 안전관리가 가능하다. 그러나, 일부 국가에서는 주소를 표기할 때 행정구역 레벨명을 생략하여 작성하고 있으며, 동일한 지명이 여러 행정구역 레벨에서 중복되는 경우가 있어 주소로부터 행정구역 레벨을 정확히 분류하는 일은 쉽지 않다. 본 연구에서는 이러한 경우에 적합한 딥러닝 기반 행정구역 레벨 분류 모델을 제안하고, 실제 해외 식품회사 주소 데이터에 대하여 검증한다. 구체적으로 다중 레이블 분류 모델에서 멱집합(Label Powerset)을 이용해 훈련하는 방식을 사용한다. 제안된 기법의 검증을 위해 식약처에 등록된 에콰도르 및 베트남에 있는 해외 제조업소 주소에 대하여 정확도를 검증하였으며, 기존의 분류 모델보다 정확도가 각각 28.1% 및 13% 정도 향상되었다.

      • KCI등재

        다중 레이블 분류를 활용한 안면 피부 질환 인식에 관한 연구

        임채현 ( Chae Hyun Lim ),손민지 ( Son Min Ji ),김명호 ( Kim Myung Ho ) 한국정보처리학회 2021 정보처리학회 논문지(KTSDE) Vol.10 No.12

        Recently, as people's interest in facial skin beauty has increased, research on skin disease recognition for facial skin beauty is being conducted by using deep learning. These studies recognized a variety of skin diseases, including acne. Existing studies can recognize only the single skin diseases, but skin diseases that occur on the face can enact in a more diverse and complex manner. Therefore, in this paper, complex skin diseases such as acne, blackheads, freckles, age spots, normal skin, and whiteheads are identified using the Inception-ResNet V2 deep learning mode with multi-label classification. The accuracy was 98.8%, hamming loss was 0.003, and precision, recall, F1-Score achieved 96.6% or more for each single class.

      • KCI등재

        다중 레이블 분류의 정확도 향상을 위한 스킵 연결 오토인코더 기반 레이블 임베딩 방법론

        김무성(Museong Kim),김남규(Namgyu Kim) 한국지능정보시스템학회 2021 지능정보연구 Vol.27 No.3

        Recently, with the development of deep learning technology, research on unstructured data analysis is being actively conducted, and it is showing remarkable results in various fields such as classification, summary, and generation. Among various text analysis fields, text classification is the most widely used technology in academia and industry. Text classification includes binary class classification with one label among two classes, multi-class classification with one label among several classes, and multi-label classification with multiple labels among several classes. In particular, multi-label classification requires a different training method from binary class classification and multi-class classification because of the characteristic of having multiple labels. In addition, since the number of labels to be predicted increases as the number of labels and classes increases, there is a limitation in that performance improvement is difficult due to an increase in prediction difficulty. To overcome these limitations, (i) compressing the initially given high-dimensional label space into a low-dimensional latent label space, (ii) after performing training to predict the compressed label, (iii) restoring the predicted label to the high-dimensional original label space, research on label embedding is being actively conducted. Typical label embedding techniques include Principal Label Space Transformation (PLST), Multi-Label Classification via Boolean Matrix Decomposition (MLC-BMaD), and Bayesian Multi-Label Compressed Sensing (BML-CS). However, since these techniques consider only the linear relationship between labels or compress the labels by random transformation, it is difficult to understand the non-linear relationship between labels, so there is a limitation in that it is not possible to create a latent label space sufficiently containing the information of the original label. Recently, there have been increasing attempts to improve performance by applying deep learning technology to label embedding. Label embedding using an autoencoder, a deep learning model that is effective for data compression and restoration, is representative. However, the traditional autoencoder-based label embedding has a limitation in that a large amount of information loss occurs when compressing a high-dimensional label space having a myriad of classes into a low-dimensional latent label space. This can be found in the gradient loss problem that occurs in the backpropagation process of learning. To solve this problem, skip connection was devised, and by adding the input of the layer to the output to prevent gradient loss during backpropagation, efficient learning is possible even when the layer is deep. Skip connection is mainly used for image feature extraction in convolutional neural networks, but studies using skip connection in autoencoder or label embedding process are still lacking. Therefore, in this study, we propose an autoencoder-based label embedding methodology in which skip connections are added to each of the encoder and decoder to form a low-dimensional latent label space that reflects the information of the high-dimensional label space well. In addition, the proposed methodology was applied to actual paper keywords to derive the high-dimensional keyword label space and the low-dimensional latent label space. Using this, we conducted an experiment to predict the compressed keyword vector existing in the latent label space from the paper abstract and to evaluate the multi-label classification by restoring the predicted keyword vector back to the original label space. As a result, the accuracy, precision, recall, and F1 score used as performance indicators showed far superior performance in multi-label classification based on the proposed methodology compared to traditional multi-label classification methods. This can be seen that the low-dimensional latent label space derived through the proposed methodology we

      • KCI등재

        특허문서의 IPC 계층별 분류기 생성: B섹션 기준 섹션부터 서브클래스까지

        박수현,김진 사)한국빅데이터학회 2025 한국빅데이터학회 학회지 Vol.10 No.1

        Since the establishment of the Korean Intellectual Property Office (KIPO), patent applications across various fields have continued to increase in Korea, and this upward trend is expected to persist. However, the patent examination process still takes more than five years. To address this issue, this paper proposes a classifier to automate IPC code prediction within the patent examination workflow, aiming to reduce examination time and improve efficiency. Given that Section B contains a large number of classes but relatively few patents compared to other sections, the scope of the proposed classifier is limited to the subclasses of Section B. Furthermore, to evaluate the impact of independent claims on classification performance, we constructed datasets both with and without independent claims. We then developed classifiers using four different techniques—Multinomial Naïve Bayes, XGBoost, LightGBM, and Random Forest— and compared their performance.

      연관 검색어 추천

      이 검색어로 많이 본 자료

      활용도 높은 자료

      해외이동버튼