머신러닝에 중요한 데이터 셋의 구성과 품질을 올리는 기술인 데이터 증강은 적은 양의 데이터를 바탕으로 다양한 알고리즘을 통해 데이터의 양을 늘리는 기술이다. 본 연구에서는 대규모 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=A108944558
2024
Korean
KCI등재
학술저널
61-69(9쪽)
0
상세조회0
다운로드머신러닝에 중요한 데이터 셋의 구성과 품질을 올리는 기술인 데이터 증강은 적은 양의 데이터를 바탕으로 다양한 알고리즘을 통해 데이터의 양을 늘리는 기술이다. 본 연구에서는 대규모 ...
머신러닝에 중요한 데이터 셋의 구성과 품질을 올리는 기술인 데이터 증강은 적은 양의 데이터를 바탕으로 다양한 알고리즘을 통해 데이터의 양을 늘리는 기술이다. 본 연구에서는 대규모 언어 모델인 GPT(Generative Pre-trained Transformer)를 활용한 데이터 증강으로 감정 분석 모델의 성능을 향상시키는 방법을 제안하고 평가한다. 데이터 셋의 클래스 불균형 문제를 해결하기 위해 가중치를 적용한 로직을 사용하였고, 생성된 데이터의 품질 및 다양성에 대한 한계를 극복하기 위해 프롬프트 엔지니어링을 적용했다. 실험결과, 제안한 방법은 데이터의 품질을 유지하면서 다양성을 높이고, 클래스 불균형 문제를 효과적으로 해결할 수 있어 KoBERT 모델을 이용해 GPT를 활용한 데이터 증강이 모델의 성능을 향상시킬 수 있음을 보였다.
다국어 초록 (Multilingual Abstract)
Data augmentation, a technology that increases the composition and quality of datasets important for machine learning, is a technology that increases the amount of data through various algorithms based on a small amount of data. In this study, we prop...
Data augmentation, a technology that increases the composition and quality of datasets important for machine learning, is a technology that increases the amount of data through various algorithms based on a small amount of data. In this study, we propose and evaluate ways to improve the performance of emotion analysis models by augmenting data using a large language model, Generative Pre-trained Transformer(GPT). We used weighted logic to solve the class imbalance problem of datasets and applied prompt engineering to overcome limitations on the quality and diversity of generated data. As a result of the experiment, it was shown that the proposed method can increase diversity while maintaining the quality of the data and effectively solve the class imbalance problem, so that data augmentation using GPT can improve the performance of the model using the KoBERT model.
참고문헌 (Reference)
1 C. Shorten, "Text Data Augmentation for Deep Learning" 8 (8): 2021
2 N. V. Chawla, "Smote : Synthetic minority oversampling technique" 16 : 2002
3 G. Parascandolo, "Recurrent neuralnetworks for polyphonic sound event detection inreal life recordings" 2016
4 이수환 ; 송기상, "Prompt engineering to improve the performance of teaching and learning materials Recommendation of Generative Artificial Intelligence" 28 (28): 195-204, 2023
5 R. Caruana, "Overfitting in neural nets : Backpropagation, conjugate gradient, and early stopping" 381-387, 2000
6 J. Wang, "Measurement of text similarity : a survey" 11 (11): 421-, 2020
7 S. Hu, "MSMOTE:improving classification performance when training data is imbalanced" 2 : 2009
8 이원민 ; 온병원, "Generating Emotional Sentences Through Sentiment and Emotion Word Masking-based BERT and GPT Pipeline Method" 19 (19): 29-40, 2021
9 AI Hub, "Emotional conversation corpus"
10 D. Baidoo-Anu, "Education in the Era of Generative Artificial Intelligence(AI) : Understanding the Potential Benefits of ChatGPT in Promoting Teaching and Learning"
1 C. Shorten, "Text Data Augmentation for Deep Learning" 8 (8): 2021
2 N. V. Chawla, "Smote : Synthetic minority oversampling technique" 16 : 2002
3 G. Parascandolo, "Recurrent neuralnetworks for polyphonic sound event detection inreal life recordings" 2016
4 이수환 ; 송기상, "Prompt engineering to improve the performance of teaching and learning materials Recommendation of Generative Artificial Intelligence" 28 (28): 195-204, 2023
5 R. Caruana, "Overfitting in neural nets : Backpropagation, conjugate gradient, and early stopping" 381-387, 2000
6 J. Wang, "Measurement of text similarity : a survey" 11 (11): 421-, 2020
7 S. Hu, "MSMOTE:improving classification performance when training data is imbalanced" 2 : 2009
8 이원민 ; 온병원, "Generating Emotional Sentences Through Sentiment and Emotion Word Masking-based BERT and GPT Pipeline Method" 19 (19): 29-40, 2021
9 AI Hub, "Emotional conversation corpus"
10 D. Baidoo-Anu, "Education in the Era of Generative Artificial Intelligence(AI) : Understanding the Potential Benefits of ChatGPT in Promoting Teaching and Learning"
11 J. Wei, "EDA : Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks"
12 N. Srivastava, "Dropout : A Simple Way to Prevent Neural Networks from Overfitting" 15 (15): 1929-1958, 2014
13 J. Salamon, "Deep convolutional neuralnetworks and dataaugmentation for environmental soundclassification" 24 (24): 2017
14 Y. LeCun, "Deep Learning" 521 : 436-444, 2015
15 T. Sorensen, "An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels"
16 A. A. Khan, "A review of ensemble learning and data augmentation models for class imbalanced problems: combination, implementation and evaluation" 244 : 2024
17 조희찬 ; 문종섭, "A layered-wise data augmenting algorithm for small sampling data" 20 (20): 65-72, 2019
18 G. E. Hinton, "A fast learning algorithm for deep belief nets" 18 (18): 2006
19 J. Ye, "A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models"
Health-AutoML: 헬스케어 데이터 분석을 위한 자동 적응형 다층 스태킹 앙상블 학습 프레임워크
ARKit을 활용한 모바일 AR 퍼포먼스 드로잉 콘텐츠 개발
Bayesian neural network를 활용한 화재 오인식 개선 연구
LiDAR 및 UWB 기반 자율주행 모드 자동 전환 기법