RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    감정분석 모델성능 향상을 위한 GPT기반 데이터 증강 방법 = GPT-based Data Augmentation Method for Improving Emotion Analysis Model Performance

    한글로보기

    https://www.riss.kr/link?id=A108944558

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    머신러닝에 중요한 데이터 셋의 구성과 품질을 올리는 기술인 데이터 증강은 적은 양의 데이터를 바탕으로 다양한 알고리즘을 통해 데이터의 양을 늘리는 기술이다. 본 연구에서는 대규모 언어 모델인 GPT(Generative Pre-trained Transformer)를 활용한 데이터 증강으로 감정 분석 모델의 성능을 향상시키는 방법을 제안하고 평가한다. 데이터 셋의 클래스 불균형 문제를 해결하기 위해 가중치를 적용한 로직을 사용하였고, 생성된 데이터의 품질 및 다양성에 대한 한계를 극복하기 위해 프롬프트 엔지니어링을 적용했다. 실험결과, 제안한 방법은 데이터의 품질을 유지하면서 다양성을 높이고, 클래스 불균형 문제를 효과적으로 해결할 수 있어 KoBERT 모델을 이용해 GPT를 활용한 데이터 증강이 모델의 성능을 향상시킬 수 있음을 보였다.
    번역하기

    머신러닝에 중요한 데이터 셋의 구성과 품질을 올리는 기술인 데이터 증강은 적은 양의 데이터를 바탕으로 다양한 알고리즘을 통해 데이터의 양을 늘리는 기술이다. 본 연구에서는 대규모 ...

    머신러닝에 중요한 데이터 셋의 구성과 품질을 올리는 기술인 데이터 증강은 적은 양의 데이터를 바탕으로 다양한 알고리즘을 통해 데이터의 양을 늘리는 기술이다. 본 연구에서는 대규모 언어 모델인 GPT(Generative Pre-trained Transformer)를 활용한 데이터 증강으로 감정 분석 모델의 성능을 향상시키는 방법을 제안하고 평가한다. 데이터 셋의 클래스 불균형 문제를 해결하기 위해 가중치를 적용한 로직을 사용하였고, 생성된 데이터의 품질 및 다양성에 대한 한계를 극복하기 위해 프롬프트 엔지니어링을 적용했다. 실험결과, 제안한 방법은 데이터의 품질을 유지하면서 다양성을 높이고, 클래스 불균형 문제를 효과적으로 해결할 수 있어 KoBERT 모델을 이용해 GPT를 활용한 데이터 증강이 모델의 성능을 향상시킬 수 있음을 보였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Data augmentation, a technology that increases the composition and quality of datasets important for machine learning, is a technology that increases the amount of data through various algorithms based on a small amount of data. In this study, we propose and evaluate ways to improve the performance of emotion analysis models by augmenting data using a large language model, Generative Pre-trained Transformer(GPT). We used weighted logic to solve the class imbalance problem of datasets and applied prompt engineering to overcome limitations on the quality and diversity of generated data. As a result of the experiment, it was shown that the proposed method can increase diversity while maintaining the quality of the data and effectively solve the class imbalance problem, so that data augmentation using GPT can improve the performance of the model using the KoBERT model.
    번역하기

    Data augmentation, a technology that increases the composition and quality of datasets important for machine learning, is a technology that increases the amount of data through various algorithms based on a small amount of data. In this study, we prop...

    Data augmentation, a technology that increases the composition and quality of datasets important for machine learning, is a technology that increases the amount of data through various algorithms based on a small amount of data. In this study, we propose and evaluate ways to improve the performance of emotion analysis models by augmenting data using a large language model, Generative Pre-trained Transformer(GPT). We used weighted logic to solve the class imbalance problem of datasets and applied prompt engineering to overcome limitations on the quality and diversity of generated data. As a result of the experiment, it was shown that the proposed method can increase diversity while maintaining the quality of the data and effectively solve the class imbalance problem, so that data augmentation using GPT can improve the performance of the model using the KoBERT model.

    더보기

    참고문헌 (Reference)

    1 C. Shorten, "Text Data Augmentation for Deep Learning" 8 (8): 2021

    2 N. V. Chawla, "Smote : Synthetic minority oversampling technique" 16 : 2002

    3 G. Parascandolo, "Recurrent neuralnetworks for polyphonic sound event detection inreal life recordings" 2016

    4 이수환 ; 송기상, "Prompt engineering to improve the performance of teaching and learning materials Recommendation of Generative Artificial Intelligence" 28 (28): 195-204, 2023

    5 R. Caruana, "Overfitting in neural nets : Backpropagation, conjugate gradient, and early stopping" 381-387, 2000

    6 J. Wang, "Measurement of text similarity : a survey" 11 (11): 421-, 2020

    7 S. Hu, "MSMOTE:improving classification performance when training data is imbalanced" 2 : 2009

    8 이원민 ; 온병원, "Generating Emotional Sentences Through Sentiment and Emotion Word Masking-based BERT and GPT Pipeline Method" 19 (19): 29-40, 2021

    9 AI Hub, "Emotional conversation corpus"

    10 D. Baidoo-Anu, "Education in the Era of Generative Artificial Intelligence(AI) : Understanding the Potential Benefits of ChatGPT in Promoting Teaching and Learning"

    1 C. Shorten, "Text Data Augmentation for Deep Learning" 8 (8): 2021

    2 N. V. Chawla, "Smote : Synthetic minority oversampling technique" 16 : 2002

    3 G. Parascandolo, "Recurrent neuralnetworks for polyphonic sound event detection inreal life recordings" 2016

    4 이수환 ; 송기상, "Prompt engineering to improve the performance of teaching and learning materials Recommendation of Generative Artificial Intelligence" 28 (28): 195-204, 2023

    5 R. Caruana, "Overfitting in neural nets : Backpropagation, conjugate gradient, and early stopping" 381-387, 2000

    6 J. Wang, "Measurement of text similarity : a survey" 11 (11): 421-, 2020

    7 S. Hu, "MSMOTE:improving classification performance when training data is imbalanced" 2 : 2009

    8 이원민 ; 온병원, "Generating Emotional Sentences Through Sentiment and Emotion Word Masking-based BERT and GPT Pipeline Method" 19 (19): 29-40, 2021

    9 AI Hub, "Emotional conversation corpus"

    10 D. Baidoo-Anu, "Education in the Era of Generative Artificial Intelligence(AI) : Understanding the Potential Benefits of ChatGPT in Promoting Teaching and Learning"

    11 J. Wei, "EDA : Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks"

    12 N. Srivastava, "Dropout : A Simple Way to Prevent Neural Networks from Overfitting" 15 (15): 1929-1958, 2014

    13 J. Salamon, "Deep convolutional neuralnetworks and dataaugmentation for environmental soundclassification" 24 (24): 2017

    14 Y. LeCun, "Deep Learning" 521 : 436-444, 2015

    15 T. Sorensen, "An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels"

    16 A. A. Khan, "A review of ensemble learning and data augmentation models for class imbalanced problems: combination, implementation and evaluation" 244 : 2024

    17 조희찬 ; 문종섭, "A layered-wise data augmenting algorithm for small sampling data" 20 (20): 65-72, 2019

    18 G. E. Hinton, "A fast learning algorithm for deep belief nets" 18 (18): 2006

    19 J. Ye, "A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models"

    더보기

    동일학술지(권/호) 다른 논문

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼