RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Synthetic Data Generation Techniques using GPT-3 Model for Enhancing Imbalanced Sentiment Analysis

    한글로보기

    https://www.riss.kr/link?id=T17054041

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In the sentiment analysis field, the prevalence of imbalanced data has presented a significant challenge to achieving accurate classification performance, especially in the case of user-generated content. Usually, most sentiment analysis datasets have tended to over-represent popular sentiments while leaving minority sentiments underrepresented, resulting in classifiers that perform well on majority classes but poorly on minority ones. Previous approaches to addressing this imbalance, such as the use of Generative Adversarial Networks (GANs) for synthetic data generation, have shown potential but also limitations, indicating a need for more advanced and effective methods. This dissertation introduces an innovative approach to overcoming this challenge by leveraging the GPT-3 model, a state-of-the-art language model known for its advanced ability to generate human-like text across a vast array of applications. Recognized for its deep learning capabilities and extensive training on diverse data, GPT-3 offers a promising solution for generating synthetic data that can balance sentiment classes effectively. This study focused on two primary objectives: first, to generate good-quality synthetic data to balance the training dataset in imbalanced sentiment analysis scenarios using two distinct strategies—fine-tuning the GPT-3 model (GPT-FT) and employing a sentence-by-sentence generation technique (GPT-SS); and second, to improve sentiment classification performance with a balanced dataset using sophisticated deep learning models, including RNN, CNN, LSTM, BiLSTM, and GRU, with GloVe embeddings for feature extraction. The methodology employed in this study involved generating synthetic data using the GPT-3 model to address the specific needs of underrepresented classes within the Coursera review dataset, a widely recognized source of user-generated content in educational sentiment analysis. This synthetic data was then evaluated for its quality based on novelty, diversity, and the relevance with the Coursera review context. The results indicated that while both GPT-FT and GPT-SS methods were effective in enhancing classification performance, GPT-FT demonstrated a better ability to generate more novel and diverse synthetic data. However, GPT-FT also faced challenges such as the occasional generation of anomalous data, highlighting areas for further refinement. Conversely, GPT-SS was able to generate synthetic data without anomalies but encountered problems with high similarity among the generated data. Additionally, the evaluation of sentiment classification showed that GPT-FT achieved higher performance compared to GPT-SS. Conclusively, this dissertation demonstrated that synthetic data augmentation using GPT-3 can significantly improve sentiment analysis performance by addressing the imbalance in class distribution. The proposed methods, GPT-FT and GPT-SS, not only contributed to the field of sentiment analysis by providing more balanced datasets but also by offering insights into the effective generation and evaluation of synthetic data. The study suggests further research into more sophisticated anomaly detection and data similarity management techniques to enhance the quality of synthetic data. Overall, the research establishes a foundational approach for using advanced language models like GPT-3 to tackle prevalent issues in sentiment analysis and opens new avenues for exploring the application of synthetic data in other areas of text analytics.
    번역하기

    In the sentiment analysis field, the prevalence of imbalanced data has presented a significant challenge to achieving accurate classification performance, especially in the case of user-generated content. Usually, most sentiment analysis datasets have...

    In the sentiment analysis field, the prevalence of imbalanced data has presented a significant challenge to achieving accurate classification performance, especially in the case of user-generated content. Usually, most sentiment analysis datasets have tended to over-represent popular sentiments while leaving minority sentiments underrepresented, resulting in classifiers that perform well on majority classes but poorly on minority ones. Previous approaches to addressing this imbalance, such as the use of Generative Adversarial Networks (GANs) for synthetic data generation, have shown potential but also limitations, indicating a need for more advanced and effective methods. This dissertation introduces an innovative approach to overcoming this challenge by leveraging the GPT-3 model, a state-of-the-art language model known for its advanced ability to generate human-like text across a vast array of applications. Recognized for its deep learning capabilities and extensive training on diverse data, GPT-3 offers a promising solution for generating synthetic data that can balance sentiment classes effectively. This study focused on two primary objectives: first, to generate good-quality synthetic data to balance the training dataset in imbalanced sentiment analysis scenarios using two distinct strategies—fine-tuning the GPT-3 model (GPT-FT) and employing a sentence-by-sentence generation technique (GPT-SS); and second, to improve sentiment classification performance with a balanced dataset using sophisticated deep learning models, including RNN, CNN, LSTM, BiLSTM, and GRU, with GloVe embeddings for feature extraction. The methodology employed in this study involved generating synthetic data using the GPT-3 model to address the specific needs of underrepresented classes within the Coursera review dataset, a widely recognized source of user-generated content in educational sentiment analysis. This synthetic data was then evaluated for its quality based on novelty, diversity, and the relevance with the Coursera review context. The results indicated that while both GPT-FT and GPT-SS methods were effective in enhancing classification performance, GPT-FT demonstrated a better ability to generate more novel and diverse synthetic data. However, GPT-FT also faced challenges such as the occasional generation of anomalous data, highlighting areas for further refinement. Conversely, GPT-SS was able to generate synthetic data without anomalies but encountered problems with high similarity among the generated data. Additionally, the evaluation of sentiment classification showed that GPT-FT achieved higher performance compared to GPT-SS. Conclusively, this dissertation demonstrated that synthetic data augmentation using GPT-3 can significantly improve sentiment analysis performance by addressing the imbalance in class distribution. The proposed methods, GPT-FT and GPT-SS, not only contributed to the field of sentiment analysis by providing more balanced datasets but also by offering insights into the effective generation and evaluation of synthetic data. The study suggests further research into more sophisticated anomaly detection and data similarity management techniques to enhance the quality of synthetic data. Overall, the research establishes a foundational approach for using advanced language models like GPT-3 to tackle prevalent issues in sentiment analysis and opens new avenues for exploring the application of synthetic data in other areas of text analytics.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    감정 분석 분야에서, 데이터의 불균형은 특히 사용자 생성 콘텐츠의 경우, 정확한 분류 성능을 달성하는 데 중대한 도전이 되었다. 일반적으로 대부분의 감정 분석 데이터셋은 인기 있는 감정을 과대표현하는 경향이 있으며 소수 감정은 상대적으로 부족하게 표현되어, 다수 클래스에 대해서는 잘 작동하지만 소수 클래스에 대해서는 성능이 떨어지는 분류기를 만들게 된다. 이러한 불균형을 해결하기 위한 이전의 접근법들, 예를 들어 합성 데이터 생성을 위한 생성적 적대 신경망(GANs)의 사용은 잠재력을 보였지만 한계도 있음을 나타내며, 더 발전된 효과적인 방법이 필요함을 시사했다. 이논문은 GPT-3 모델을 활용하여 이러한 도전을 극복하는 혁신적인 접근법을 소개한다. GPT-3는 다양한 애플리케이션에서 인간과 같은 텍스트를 생성할 수 있는 탁월한 능력으로 알려진 최신 언어 모델이다. 깊은 학습 능력과 다양한 데이터에 대한 광범위한 훈련을 통해 인정받은 GPT-3는 감정 클래스를 효과적으로 균형 잡히게 할 수 있는 합성 데이터를 생성하는 유망한 해결책을 제공하였다. 이 연구는 두 가지 주요 목표에 중점을 두었다. 첫째, 불균형 감정 분석 시나리오에서 훈련 데이터셋을 균형잡히게 하기 위해 GPT-3 모델을 미세 조정하는 방법(GPT-FT)과 문장별 생성 기술(GPT-SS)을 사용하여 양질의 합성 데이터를 생성하는 것이다. 둘째, GloVe 임베딩을 특징 추출에 사용하면서 RNN, CNN, LSTM, BiLSTM, GRU 등 고급 딥러닝 모델을 사용하여 균형 잡힌 데이터셋으로 감정 분류 성능을 향상시키는 것이다. 이 연구에서 사용한 방법론은 교육 감정 분석에서 사용자 생성 콘텐츠의 널리 알려진 출처인 Coursera 리뷰 데이터셋 내 소수 클래스의 특정 요구를 해결하기 위해 GPT-3 모델을 사용하여 합성 데이터를 생성하는 것을 포함했다. 이 합성 데이터는 그 후 신규성, 다양성 및 일관성을 기준으로 그 품질이 평가되었다. 결과는 GPT-FT와 GPT-SS 방법 모두 분류 성능을 향상시키는 데 효과적이었지만, GPT-FT가 더욱 새롭고 다양한 합성 데이터를 생성하는 능력이 더 뛰어났음을 보여주었다. 그러나 때때로 이상 데이터 생성과 같은 문제도 발생하여 추가 개선이 필요한 영역을 강조했다. 반면에 GPT-SS는 이상 데이터 없이 합성 데이터를 생성할 수 있었지만 생성된 데이터 간에 높은 유사성을 보이는 문제가 있었다. 또한 감정 분류 평가는 GPT-FT가 GPT-SS에 비해 더 높은 성능을 달성했음을 보여주었다. 결론적으로, 이 논문은 GPT-3를 사용한 합성 데이터 증강이 클래스 분포의 불균형을 해결함으로써 감정 분석 성능을 크게 향상시킬 수 있음을 입증했다. 제안된 방법들, GPT-FT 및 GPT-SS는 감정 분석 분야에 더 균형 잡힌 데이터셋을 제공함으로써 기여할 뿐만 아니라, 합성 데이터의 효과적인 생성 및 평가에 대한 통찰을 제공한다. 연구는 합성 데이터의 품질을 향상시키기 위해 더 정교한 이상 감지 및 데이터 유사성 관리 기술에 대한 추가 연구를 제안한다. 전반적으로, 이 연구는 감정 분석에서 널리 있는 문제를 해결하기 위해 GPT-3와 같은 고급 언어 모델을 사용하는 기초적인 접근 방식을 확립하고, 다른 텍스트 분석 분야에서 합성 데이터의 활용을 탐구할 새로운 길을 열었다.
    번역하기

    감정 분석 분야에서, 데이터의 불균형은 특히 사용자 생성 콘텐츠의 경우, 정확한 분류 성능을 달성하는 데 중대한 도전이 되었다. 일반적으로 대부분의 감정 분석 데이터셋은 인기 있는 감...

    감정 분석 분야에서, 데이터의 불균형은 특히 사용자 생성 콘텐츠의 경우, 정확한 분류 성능을 달성하는 데 중대한 도전이 되었다. 일반적으로 대부분의 감정 분석 데이터셋은 인기 있는 감정을 과대표현하는 경향이 있으며 소수 감정은 상대적으로 부족하게 표현되어, 다수 클래스에 대해서는 잘 작동하지만 소수 클래스에 대해서는 성능이 떨어지는 분류기를 만들게 된다. 이러한 불균형을 해결하기 위한 이전의 접근법들, 예를 들어 합성 데이터 생성을 위한 생성적 적대 신경망(GANs)의 사용은 잠재력을 보였지만 한계도 있음을 나타내며, 더 발전된 효과적인 방법이 필요함을 시사했다. 이논문은 GPT-3 모델을 활용하여 이러한 도전을 극복하는 혁신적인 접근법을 소개한다. GPT-3는 다양한 애플리케이션에서 인간과 같은 텍스트를 생성할 수 있는 탁월한 능력으로 알려진 최신 언어 모델이다. 깊은 학습 능력과 다양한 데이터에 대한 광범위한 훈련을 통해 인정받은 GPT-3는 감정 클래스를 효과적으로 균형 잡히게 할 수 있는 합성 데이터를 생성하는 유망한 해결책을 제공하였다. 이 연구는 두 가지 주요 목표에 중점을 두었다. 첫째, 불균형 감정 분석 시나리오에서 훈련 데이터셋을 균형잡히게 하기 위해 GPT-3 모델을 미세 조정하는 방법(GPT-FT)과 문장별 생성 기술(GPT-SS)을 사용하여 양질의 합성 데이터를 생성하는 것이다. 둘째, GloVe 임베딩을 특징 추출에 사용하면서 RNN, CNN, LSTM, BiLSTM, GRU 등 고급 딥러닝 모델을 사용하여 균형 잡힌 데이터셋으로 감정 분류 성능을 향상시키는 것이다. 이 연구에서 사용한 방법론은 교육 감정 분석에서 사용자 생성 콘텐츠의 널리 알려진 출처인 Coursera 리뷰 데이터셋 내 소수 클래스의 특정 요구를 해결하기 위해 GPT-3 모델을 사용하여 합성 데이터를 생성하는 것을 포함했다. 이 합성 데이터는 그 후 신규성, 다양성 및 일관성을 기준으로 그 품질이 평가되었다. 결과는 GPT-FT와 GPT-SS 방법 모두 분류 성능을 향상시키는 데 효과적이었지만, GPT-FT가 더욱 새롭고 다양한 합성 데이터를 생성하는 능력이 더 뛰어났음을 보여주었다. 그러나 때때로 이상 데이터 생성과 같은 문제도 발생하여 추가 개선이 필요한 영역을 강조했다. 반면에 GPT-SS는 이상 데이터 없이 합성 데이터를 생성할 수 있었지만 생성된 데이터 간에 높은 유사성을 보이는 문제가 있었다. 또한 감정 분류 평가는 GPT-FT가 GPT-SS에 비해 더 높은 성능을 달성했음을 보여주었다. 결론적으로, 이 논문은 GPT-3를 사용한 합성 데이터 증강이 클래스 분포의 불균형을 해결함으로써 감정 분석 성능을 크게 향상시킬 수 있음을 입증했다. 제안된 방법들, GPT-FT 및 GPT-SS는 감정 분석 분야에 더 균형 잡힌 데이터셋을 제공함으로써 기여할 뿐만 아니라, 합성 데이터의 효과적인 생성 및 평가에 대한 통찰을 제공한다. 연구는 합성 데이터의 품질을 향상시키기 위해 더 정교한 이상 감지 및 데이터 유사성 관리 기술에 대한 추가 연구를 제안한다. 전반적으로, 이 연구는 감정 분석에서 널리 있는 문제를 해결하기 위해 GPT-3와 같은 고급 언어 모델을 사용하는 기초적인 접근 방식을 확립하고, 다른 텍스트 분석 분야에서 합성 데이터의 활용을 탐구할 새로운 길을 열었다.

    더보기

    목차 (Table of Contents)

    • I.Introduction 1
    • A.Background 1
    • B.Problem Formulation 3
    • C.Research Objective 4
    • D.Organization 5
    • I.Introduction 1
    • A.Background 1
    • B.Problem Formulation 3
    • C.Research Objective 4
    • D.Organization 5
    • II.Related Works 6
    • A.Text Generation 6
    • B.Imbalanced Sentiment Analysis 8
    • III.Sentiment Analysis with GPT-3 based Synthetic Data Generation 10
    • A.Imbalanced Dataset of Coursera Review 10
    • B.The Proposed Approach for Enhancing ImbalancedSentiment Analysis 11
    • 1.GPT-FT: Fine-tuning GPT-3 based Synthetic Data Generation 13
    • 2.GPT-SS: GPT-3 Sentence-by-Sentence Synthetic DataGeneration 23
    • C.Sentiment Classification using Deep Learning Models 31
    • IV.Evaluation of The Proposed Approach 39
    • A.The Evaluation of GPT-FT for Coursera Review Dataset 39
    • 1.GPT-FT Generated Synthetic Data 39
    • 2.Evaluation of GPT-FT Generated Synthetic Data 40
    • B.The evaluation of GPT-SS for Coursera Review Dataset 46
    • 1.GPT-SS Generated Synthetic Data 46
    • 2.Evaluation of GPT-SS Generated Synthetic Data 49
    • C.The Evaluation of Sentiment Classification 57
    • 1.Distribution of Dataset for Sentiment Classification 57
    • 2.Sentiment Classification Results 59
    • D.Discussion: GPT-FT and GPT-SS Comparison 63
    • V.Conclusion and Future Work 70
    • Bibliography 74
    • Abstract (in Korean) 83
    더보기

    참고문헌 (Reference)

    1. Attention Is All You Need., A. Vaswani et al., arXiv 2023 AccessedOnline Available http//arxiv org/abs/1706.03762, , 2024

    2. Language Models are Few-Shot Learners., T. B. Brown et al., arXiv Jul. 22,2020 AccessedOnline Availablehttp//arxiv org/abs/2005.14165, , 2024

    3. Can GPT-3 Pass a Writer’s Turing Test?, J. Chun, K. Elkins and, vol. 5, no. 2, doi10.22148/001c.17212, , 2020

    4. Not Enough Data? Deep Learning to the Rescue, A. Anaby-Tavor et al., arXiv Accessed 2024Online Availablehttp//arxiv org/abs/1911.03118, , 2019

    5. The class imbalance problemA systematic study1, N. Japkowicz and, S. Stephen, Intell. Data Anal., vol. 6, no. 5, pp. 429–449, doi10.3233/IDA-2002-6504, , 2002

    6. GPT-3Its Nature, Scope, Limits, andConsequences, M. Chiriatti, L. Floridi and, Minds Mach., vol. 30, no. 4, pp. 681–694, doi:10.1007/s11023-020-09548-1, , 2020

    7. Data Imbalance Problem in Text Classification,in, G. Sun and, Y. Li, Y. Zhu, 2010 Third International Symposium on Information Processing, Qingdao,Shandong, ChinaIEEE, pp. 301–305. doi10.1109/ISIP.2010.47, , 2010

    8. The surveyText generation models in deeplearning, T. Iqbal and, S. Qureshi, vol. 34, no. 6, pp. 2515–2528, doi10.1016/j. jksuci.2020.04.001, , 2022

    9. 53 Aaron Courville - Deep Learningpre-pub version, Yoshua Bengio, Ian Goodfellow, -MIT Press pdf, , 2016

    10. Dealing with Data Imbalance in TextClassification, C. Padurariu and, M. E. Breaban, vol. 159, pp. 736–745, doi:10.1016/j. procs.2019.09.229, , 2019

    1. Attention Is All You Need., A. Vaswani et al., arXiv 2023 AccessedOnline Available http//arxiv org/abs/1706.03762, , 2024

    2. Language Models are Few-Shot Learners., T. B. Brown et al., arXiv Jul. 22,2020 AccessedOnline Availablehttp//arxiv org/abs/2005.14165, , 2024

    3. Can GPT-3 Pass a Writer’s Turing Test?, J. Chun, K. Elkins and, vol. 5, no. 2, doi10.22148/001c.17212, , 2020

    4. Not Enough Data? Deep Learning to the Rescue, A. Anaby-Tavor et al., arXiv Accessed 2024Online Availablehttp//arxiv org/abs/1911.03118, , 2019

    5. The class imbalance problemA systematic study1, N. Japkowicz and, S. Stephen, Intell. Data Anal., vol. 6, no. 5, pp. 429–449, doi10.3233/IDA-2002-6504, , 2002

    6. GPT-3Its Nature, Scope, Limits, andConsequences, M. Chiriatti, L. Floridi and, Minds Mach., vol. 30, no. 4, pp. 681–694, doi:10.1007/s11023-020-09548-1, , 2020

    7. Data Imbalance Problem in Text Classification,in, G. Sun and, Y. Li, Y. Zhu, 2010 Third International Symposium on Information Processing, Qingdao,Shandong, ChinaIEEE, pp. 301–305. doi10.1109/ISIP.2010.47, , 2010

    8. The surveyText generation models in deeplearning, T. Iqbal and, S. Qureshi, vol. 34, no. 6, pp. 2515–2528, doi10.1016/j. jksuci.2020.04.001, , 2022

    9. 53 Aaron Courville - Deep Learningpre-pub version, Yoshua Bengio, Ian Goodfellow, -MIT Press pdf, , 2016

    10. Dealing with Data Imbalance in TextClassification, C. Padurariu and, M. E. Breaban, vol. 159, pp. 736–745, doi:10.1016/j. procs.2019.09.229, , 2019

    11. Glove: Global vectors for word representation, in, J. Pennington, C. Manning, R. Socher and, Proceedings of the 2014 Conference on EmpiricalMethods in Natural Language Processing (EMNLP), Doha, Qatar:Association for Computational Linguistics pp. 1532–1543. doi:10.3115/v1/D14-1162, , 2014

    12. Text Generation forImbalanced Text Classification, S. Sinthupinyo, P. Kachamas and, S. Akkaradamrongrat, 2019 16th International Joint Conference on Computer Science and Software Engineering (JCSSE), Chonburi, ThailandIEEE, pp. 181–186. doi10.1109/JCSSE.2019.8864181, , 2019

    13. Metrics for Multi-Class Classification:an Overview., M. Grandini, G. Visani, E. Bagli and, arXiv, AccessedOnline Available http//arxiv org/abs/2008.05756, , 2020

    14. Train Test Validation SplitHow To & Best Practices., P. Baheti, AccessedOnline Available https//www. v7labs. com/blog/train-validation-test-set#h1, , 2024

    15. A Brief History of DeepLearning-Based Text Generation, A. Bas, I. Van Heerden, M. O. Topal, C. Duman and, in 2022 International Conference onComputer and Applications (ICCA), Cairo, EgyptIEEE, pp. 1–4. doi10.1109/ICCA56443.2022.10039545, , 2022

    16. Deduplicating Training Data Makes Language Models Better., K. Lee et al., arXiv AccessedOnline Availablehttp//arxiv org/abs/2107.06499, , 2022

    17. A Brief Survey of Word Embedding and Its RecentDevelopment, Q. Jiao and, S. Zhang, in 2021 IEEE 5th Advanced Information Technology,Electronic and Automation Control Conference (IAEAC), Chongqing, China:IEEE, pp. 1697–1701. doi10.1109/IAEAC50856.2021.9390956, , 2021

    18. Sentiment Classification Using ConvolutionalNeural Networks, H. Kim and, Y.-S. Jeong, vol. 9, no. 11, p. 2347, doi:10.3390/app9112347, , 2019

    19. Sentiment analysismining opinions, sentiments, and emotions, B. Liu, NewYork, NYCambridge University Press, , 2015

    20. Document Processing:Methods for Semantic Text Similarity Analysis, A. P. Johnson, V. Holmes and, A. W. Qurashi, in 2020 InternationalConference on INnovations in Intelligent SysTems and Applications (INISTA),Novi Sad, SerbiaIEEE, pp. 1–6. doi:10.1109/INISTA49547.2020.9194665., , 2020

    21. Quantifying the Effects of TextDuplication on Semantic Models, in, D. Mimno, L. Thompson and, A. Schofield, Proceedings of the 2017 Conference onEmpirical Methods in Natural Language Processing, Copenhagen, Denmark:Association for Computational Linguistics, pp. 2737–2747. doi:10.18653/v1/D17-1290, , 2017

    22. SentiGANGenerating Sentimental Texts via MixtureAdversarial Networks, K. Wang and, X. Wan, in Proceedings of the Twenty-Seventh InternationalJoint Conference on Artificial Intelligence, Stockholm, SwedenInternationalJoint Conferences on Artificial Intelligence Organization, pp.4446–4452. doi10.24963/ijcai.2018/618, , 2018

    23. Tweet Sentiment Analysis of the 2020 U. S. Presidential Election, in, H. Yue and, E. Xia, H. Liu, Companion Proceedings of the Web Conference2021, Ljubljana SloveniaACM, pp. 367–371. doi:10.1145/3442442.3452322, , 2021

    24. Deduplicating Training DataMitigates Privacy Risks in Language Models., N. Kandpal, E. Wallace and, C. Raffel, arXiv, Dec. 20, AccessedDec. 11, 2023.Online] Availablehttp://arxiv. org/abs/2202.06539, , 2022

    25. RoBERTa-GRUA Hybrid DeepLearning Model for Enhanced Sentiment Analysis, K. L. Tan, K. M. Lim, C. P. Lee and, vol. 13, no. 6,p.3915, Mar. 2023, doi10.3390/app13063915, , 2023

    26. Applications and Challenges of Sentiment Analysisin Real-life Scenarios, D. Kanojia and, A. Joshi, 2023, doi10.48550/ARXIV.2301.09912, , 2023

    27. A Survey of Sentiment Analysis:Approaches, Datasets, and Future Research,, C. P. Lee and, K. L. Tan, K. M. Lim, vol. 13, no. 7, p.4550,, doi10.3390/app13074550, , 2023

    28. SAKPA Korean Sentiment Analysis Model viaKnowledge Base and Prompt Tuning, Z. Zhang, H. Wen and, in 2023 IEEE 3rd InternationalConference on Computer Communication and Artificial Intelligence (CCAI),Taiyuan, ChinaIEEE, pp. 147–152. doi:10.1109/CCAI57533.2023.10201257, , 2023

    29. Sentiment Analysis of Imbalanced Comment TextsUnder the Framework of BiLSTM, J. Zhao, H. Wen and, in 2023 6th International Conference onArtificial Intelligence and Big Data (ICAIBD), Chengdu, ChinaIEEE, pp. 312–319. doi10.1109/ICAIBD57115.2023.10206154, , 2023

    30. Sentimentanalysis to support business decision-making. A bibliometric study, P. R. Palos-Sanchez and, R. Pozo-Barajas, J. A. Aguilar-Moreno, AIMSMath., vol. 9, no. 2, pp. 4337–4375, 2024, doi10.3934/math.2024215, , 2024

    31. Sentiment analysisA survey on designframework, applications and future scopes, M. Bordoloi and, S. K. Biswas, vol. 56, no. 11,pp. 12505–12560,, doi10.1007/s10462-023-10442-2, , 2023

    32. Can ChatGPT Understand Too? A Comparative Study on ChatGPT and Fine-tuned BERT, B. Du and, L. Ding, D. Tao, Q. Zhong, J. Liu, 2023 doi10.48550/ARXIV.2302.10198, , 2023

    33. ASystematic Literature Review on Text Generation Using Deep NeuralNetwork Models, Z. Kastrati, A. Soomro, A. S. Imran, S. M. Daudpota and, N. Fatima, IEEE Access, vol. 10, pp. 53490–53503, 2022, doi:10.1109/ACCESS.2022.3174108, , 2022

    34. A systematic study of the class imbalance problem in convolutional neural networks,, A. Maki and, M. Buda, M. A. Mazurowski, Neural Netw., vol. 106,pp. 249–259, doi10.1016/j. neunet.2018.07.011, , 2018

    35. The impactof synthetic text generation for sentiment analysis using GAN based models, A. S. Imran, S. Shaikh, R. Yang, S. M. Daudpota and, Z. Kastrati, Egypt. Inform. J., vol. 23, no. 3, pp. 547–557, doi:10.1016/j. eij.2022.05.006, , 2022

    36. Sentiment Analysis of before and after Elections:Twitter Data of U. S. Election 2020,, H. N. Chaudhry et al., Electronics, vol. 10, no. 17, p. 2082, doi10.3390/electronics10172082, , 2021

    37. Table Caption Generation in ScholarlyDocuments Leveraging Pre-trained Language Models., J. H. Xu, M. P. Kato, K. Shinden and, arXiv, Aug. 18, 2021. AccessedApr. 17, 2024Online Available http://arxiv. org/abs/2108.08111, , 2021

    38. Imbalanced Text Sentiment Classification Based onMulti-Channel BLTCN-BLSTM Self-Attention, X. Zhang, T. Cai and, vol. 23, no. 4, p.2257,, doi10.3390/s23042257, , 2023

    39. A ComprehensiveOverview and Comparative Analysis on Deep Learning ModelsCNN, RNN,LSTM, GRU., R. Mohamed, T. Perumal, N. Mustapha and, F. M. Shiri, 56 arXiv, 2023. AccessedOnline Available http//arxiv. org/abs/2305.17473, , 2024

    40. Sentiment analysis ofimbalanced datasets using BERT and ensemble stacking for deep learning, H. Nouri, L. Hassouni, H. Anoun and, N. Habbat, vol. 126, p. 106999,, doi:10.1016/j. engappai.2023.106999, , 2023

    41. Mitigating Class Imbalance in SentimentAnalysis through GPT-3-Generated Synthetic Sentences,, C. Suhaeni and, H.-S. Yong, vol. 13,no. 17, p. 9766,, doi10.3390/app13179766, , 2023

    42. Enhancing Imbalanced Sentiment AnalysisAGPT-3-Based Sentence-by-Sentence Generation Approach,, C. Suhaeni and, H.-S. Yong, vol.14, no. 2, p. 622,, doi10.3390/app14020622, , 2024

    43. Impact of Sentiment Analysis for the2020 U. S. Presidential Election on Social Media Data, in, M. Qorib, R. S. Gizaw and, J. Kim, Proceedings of the2023 8th International Conference on Machine Learning Technologies,Stockholm SwedenACM,, pp. 28–34. doi:10.1145/3589883.3589888, , 2023

    44. Prediction and analysis of IndonesiaPresidential election from Twitter using sentiment analysis, M. Meiliana, W. Budiharto and, J. Big Data, vol.5, no. 1, p. 51, doi10.1186/s40537-018-0164-1, , 2018

    45. Using AraGPT andensemble deep learning model for sentiment analysis on Arabic imbalanceddataset, H. Nouri, H. Anoun and, N. Habbat, L. Hassouni, ITM Web Conf., vol. 52, p. 02008, 2023, doi:10.1051/itmconf/20235202008, , 2008

    46. Switch-GPTAn Effective Methodfor Constrained Text Generation under Few-Shot Settings (Student Abstract), G. Shen and, C. Ma, S. Zhang, Z. Deng, vol. 36, no. 11, pp. 13011–13012,doi10.1609/aaai. v36i11.21642., , 2022

    47. Sentiment analysis of COVID-19 tweets from selected hashtags in Nigeriausing VADER and Text Blob analyser, A. Abayomi-Alli, S. Misra and, O. Abiola, O. Abayomi-Alli, O. A. Tale, J. Electr. Syst. Inf. Technol., vol. 10, no. 1, p. 5, , 2023

    48. 44 Performance Evaluation of Sentiment Analysis on Balanced andImbalanced Dataset Using Ensemble Approach,, Avinashilingam, V. Srividhya, Department, S. George and, Coimbatore, Computer Science for Home Scienceand Higher Education for Women India vol.15, no. 17, pp. 790–797, doi10.17485/IJST/v15i17.2339., , 2022

    49. Towards ImprovedClassification Accuracy on Highly Imbalanced Text Dataset Using DeepNeural Language Models, S. M. Daudpota, A. S. Imran and, Z. Kastrati, S. Shaikh, vol. 11, no. 2, p. 869, doi:10.3390/app11020869, , 2021

    50. CrowdDecision MakingSparse Representation Guided by Sentiment Analysis forLeveraging the Wisdom of the Crowd, C. Zuheros, E. Martinez-Camara, F. Herrera, E. Herrera-Viedma and, IEEE Trans. Syst. Man Cybern. Syst.,vol. 53, no. 1, pp. 369–379, Jan. 2023, doi10.1109/TSMC.2022.3180938., , 2023

    51. CatGANCategory-Aware GenerativeAdversarial Networks with Hierarchical Evolutionary Learning for CategoryText Generation, J. Wang and, Z. Liu, Z. Liang, vol. 34, no. 05, pp. 8425–8432, doi10.1609/aaai. v34i05.6361, , 2020

    52. Comparison between Machine Learning and Deep LearningApproaches for the Detection of Toxic Comments on Social Networks,, A. Bonetti, J. C. Torres, J. Vila-Francés, J. M. Vega, S. Pellerin and, M. Martínez-Sober, vol. 13, no. 10, p. 6038,, doi10.3390/app13106038., , 2023

    53. Application of Generative Adversarial Networks andShapley Algorithm Based on Easy Data Augmentation for Imbalanced TextData, J.-L. Wu and, S. Huang, vol. 12, no. 21, p. 10964, doi:10.3390/app122110964, , 2022

    54. Sentiment Analysis of Customers’ Reviews Using a HybridEvolutionary SVM-Based Approach in an Imbalanced Data Distribution, R. Obiedat et al., IEEE Access, vol. 10, pp. 22260–22273, 2022, doi:10.1109/ACCESS.2022.3149482, , 2022

    55. Training neural network classifiers for medical decision making:The effects of imbalanced datasets on classification performance,, M. A. Mazurowski, J. Y. Lo, J. M. Zurada, J. A. Baker and, G. D. Tourassi, P. A. Habas, NeuralNetw., vol. 21, no. 2–3, pp. 427–436, , 2008

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼