
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
김진수 한국산학기술학회 2025 한국산학기술학회논문지 Vol.26 No.7
최근 딥러닝 분야에서 LLM 기반의 자연어 처리에서 사전 학습과 미세 조정의 학습을 통해 다양한 분야에서 기대 이상의 높은 성능을 보여준다. 프로젝트 진행 중 발생하는 갈등은 주로 팀원 간의 의사소통 부족, 역할 및 책임의 불분명, 기대치의 차이, 작업 스타일의 충돌, 그리고 시간 관리 등에서 발생한다. 본 연구는 팀 프로젝트 수행 중 발생할 수 있는 갈등을 사전에 예측하기 위해 문장에서 감정 및 감성 정보를 추출하고, 문장 간 의미 유사도를 분석하여 갈등 여부를 판단하는 방법을 제안한다. 감정 분석은 KoBERT 기반 모델을 통해 긍정/부정 감정을 예측하고 문장 간의 문맥적 의미를 고려하여 정확도를 높인다. 문장 간 유사도 분석은 SBERT 기반 모델을 활용하여 문장을 임베딩한 후 코사인 유사도를 통해 의미적 거리까지 반영한다. 제안한 감정 및 감성 정보, 그리고 문장 간의 유사도를 고려하여 갈등을 예측한 KoBERT_EmoSentSim이 감정 및 감성만을 고려한 KoBERT_EmoSent와 문장 유사도만을 고려한 KoBERT_Sim과 비교했을 때, 정확도는 2.22%, 28.89%, F1은 2.41%, 28.04% 각각 성능이 높아짐을 보였다. 따라서 팀 내 커뮤니케이션의 정서적 흐름과 의미적 맥락을 반영하여 갈등을 조기에 예측하고, 원활한 협업 환경을 조성하는 데 기여할 수 있다. In the field of deep learning, LLM-based natural language processing has recently shown higher than expected performance in various fields through pre-learning and fine-tuning. Conflicts that occur during a project's progress mainly occur due to a lack of communication among team members, unclear roles and responsibilities, differences in expectations, conflicts in work styles, and poor time management. This study proposes the KoBERT_EmoSentSim method to extract emotion and sentiment information from sentences to analyze semantic similarity among them and to predict conflicts that might occur during team project execution. Sentiment analysis predicts positive/negative emotions using a KoBERT-based model, and it increases accuracy by considering contextual meaning between sentences. Inter-sentence similarity analysis uses an SBERT-based model to embed sentences and reflect semantic distances through cosine similarity. When KoBERT_EmoSentSim, which predicts conflicts by considering emotion and sentiment information and the similarity between sentences, was compared with KoBERT_EmoSent (which only considers emotion and sentiment) and KoBERT_Sim (which only considers sentence similarity), the accuracy was 2.22% and 28.89% higher, respectively, and F1-scores were 2.41% and 28.04% higher. Therefore, KoBERT_EmoSentSim can help predict conflicts early, creating a smooth collaborative environment by reflecting the emotional flow and semantic context of communication within a team.
KoBERT를 이용한 기업관련 신문기사 감성 분류 연구7)
현지원,이준일,조현권 한국회계학회 2022 회계학연구 Vol.47 No.4
This study explores the accuracy level of the sentiment analysis of news article sentences from Korean newspaper, using KoBERT which is a modified version of BERT developed by Google. For comparison, we use MBERT which is the multilingual version of BERT, Google Sentiment Analysis provided through Google API, and dictionary based approach. This paper finds that the accuracy level of the sentiment classification based on KoBERT is the highest at 85.7%, achieving a significantly higher level of accuracy compared to the other three models. MBERT shows the next highest accuracy level at 77.5% and the other two models provide even lower accuracy rate. We further investigate whether the sentiment classification results obtained from these four models could predict future stock return. Using cumulative future stock returns for 3 or 5 days after the news on corporation publishes, we find that the sentiment score based on the sentiment classification from the KoBERT model predicts future return better than the other three models. Overall, these findings would serve as a reference for conducting further studies related to sentiment analysis on accounting and financial text. 이 연구에서는 Google에서 개발한 BERT에 기반한 KoBERT 모형을 사용하여 한국 신문기사의 감성분석 정확도를 테스트하였다. 비교를 위해, Google에서 다국어용으로 제시한 MBERT, Google에서 API를 통해 제공하는 Google Sentiment Analysis, 그리고 사전적 접근법 을 통한 감성분석 결과를 사용하였다. 감성분석 학습 결과, KoBERT를 사용한 경우가 85.7%의 정확도를 보여, 여타 모형의 정확도에 비해 상당히 높은 수준의 정확도를 달성하는 것을 확인하였다. 다른 모형의 경우, MBERT가 77.3%의 정확도로 KoBERT에 비해 상당히 낮은 결과를 보였으며 Google Sentiment Analysis와 사전적 접근법은 더욱 낮은 정확도를 보였다. 감성분석 결과가 실질적으로 의미있는 유용한 정보를 제공하는지 확인하기 위하여 뉴스가 나온 날짜를 기준으로 3일 후, 그리고 5일 후까지 주가수익률을 종속변수로 하여 회귀분석한 결과, KoBERT를 사용한 결과가 다른 결과에 비해 미래 수익률을 더욱 잘 예측하는 것을 발견하였다. 이와 같은 결과는 추후 회계⋅재무 분야의 텍스트에 대한 감성분석을 이용한 다양한 연구를 수행하는 데 참고가 될 것으로 기대한다.
KoBERT 기반 한국어 음식점 리뷰 속성 기반 다중 레이블 감성분석
이성학,방진숙 한국정보통신학회 2025 한국정보통신학회논문지 Vol.29 No.11
온라인 음식점 리뷰가 다양해지면서 소비자는 속성별 상세 의견을 필요로 하나, 기존 감성 분석은 문서 전체의 긍/부정만을 제공한다. 본 연구는 한국어 리뷰에 특화된 속성 기반 감성 분석(ABSA) 프레임워크를 제안한다. 먼저, 총 43만 건의 네이버 지도 리뷰를 수집 및 전처리(특수문자와 이모지 제거, ET5 기반 맞춤법 교정)하고, FastText 확장법과 수작업 주석으로 ‘맛, 가격, 서비스, 분위기’ 4개 속성의 감성 사전을 구축하였다. 이어 K3 리뷰 데이터셋의 중립 제거 및 오버샘플링으로 KoBERT를 리뷰 도메인에 파인튜닝하고, monologg/kobert 기반 MultiTaskBert로 4속성×3극성의 멀티라벨·멀티클래스 분류를 구현하였다. 동일 프로토콜에서 mBERT, CNN-LSTM과 비교한 결과, 제안 모델(KoBERT)은 Exact Match 92.8%, Hamming Loss 0.021, Micro-F1 0.980, Macro-F1 0.906으로 가장 높은 성능을 보였으며, mBERT(EM 90.3%, HL 0.028, Macro-F1 0.881)와 CNN-LSTM(EM 81.2%, HL 0.056, Macro-F1 0.765)을 유의하게 상회하였다. ROC/혼동행렬 분석에서도 특히 부정 클래스 분별력이 향상되었고, 주요 오류는 약한 부정이 긍정·중립으로 해석되는 경계 사례에 집중됨을 확인하였다. 본 연구는 대규모 한국어 음식점 ABSA 데이터셋과 강력한 멀티라벨 베이스라인을 제시함으로써, 개인화 추천 및 후속 한국어 ABSA 연구의 기반을 마련한다. Consumers increasingly require aspect-level insight that document-level sentiment cannot provide. We present an ABSA framework tailored to Korean restaurant reviews. We collect and preprocess 430k Naver Maps reviews, removing special characters and emojis and applying ET5-based spelling correction. Using manual seeds and FastText expansion, we build lexicons for four aspects—taste, price, service, and atmosphere. For domain adaptation, KoBERT is fine-tuned on the KR3 set after removing neutral samples and oversampling negatives. On monologg/KoBERT, a MultiTaskBERT with a shared encoder and four parallel heads performs multi-label, multi-class classification (4 aspects × 3 polarities). Under the same protocol, KoBERT achieves Exact Match 92.8%, Hamming Loss 0.021, Micro-F1 0.980, and Macro-F1 0.906, outperforming mBERT (EM 90.3%, HL 0.028, Macro-F1 0.881) and CNN-LSTM (EM 81.2%, HL 0.056, Macro-F1 0.765). ROC and confusion-matrix analyses show stronger discrimination for negatives; remaining errors arise in borderline cases with weak negative cues alongside positive context. We release the dataset and a multi-label baseline for recommendation and ABSA research.
KoBERT 모델을 활용한 텍스트 기반 리뷰 감정 분류
정창현(Chang-Hyeon Jeong),설성중(Sung-Joong Seol),이재혁(Jae-Hyuk Lee),임지후(Ji-Hoo Lim),염찬욱(Chan-Uk Yeom),곽근창(Keun-Chang Kwa) 한국정보기술학회 2024 Proceedings of KIIT Conference Vol.2024 No.11
본 논문은 KoBERT 모델을 활용하여 한국어 텍스트에서 기쁨, 분노, 슬픔, 불안, 당황, 상처의 6가지 감정을 분류하는 모델을 제안한다. 첫 번째 단계는 KoBERT 모델을 말뭉치 데이터셋으로 파인튜닝 하는 과정이다. 두 번째 단계는 하이퍼파라미터 조정을 통해 모델 성능을 향상시키는 과정이며, 마지막 단계는 웹 크롤링으로 확보한 실제 리뷰 데이터를 사용하여 성능을 평가 및 검증하는 과정이다. 이를 위해 말뭉치 데이터셋을 활용하여 KoBERT 모델을 파인튜닝하고, 실제 리뷰 데이터를 통해 성능을 검증하였다. KoBERT 기반 감정 분류 모델은 하이퍼파라미터 최적화를 통해 Test Accuracy 0.67, F1-Score 0.69의 성능을 기록하며, 기존 모델인 LSTM, GRU 보다 성능이 향상되었다. 또한, 텍스트 데이터를 정확히 반영하여 감정 분류의 신뢰성을 높였다. 실제 영화 관람평을 입력값으로 테스트한 결과, 감정 분류 성능이 우수하여 실생활 응용 가능성도 확인되었다. This paper proposes a model for classifying six emotions—joy, anger, sadness, anxiety, embarrassment, and hurt—in Korean text using the KoBERT model. The first step involves fine-tuning the KoBERT model with a corpus dataset. The second step focuses on enhancing the model's performance by adjusting hyperparameters, and the final step involves evaluating and validating performance using actual review data collected via web crawling. For this purpose, the KoBERT model was fine-tuned with a corpus dataset, and its performance was validated with real review data. The KoBERT-based emotion classification model recorded a Test Accuracy of 0.67 and an F1-Score of 0.69 through hyperparameter optimization, outperforming previous models such as LSTM and GRU. Furthermore, it demonstrated high reliability in emotion classification by accurately reflecting textual data. Testing with actual movie reviews as input confirmed the model’s excellent emotion classification performance and practical applicability.
방승주,박요한,김지은,이공주 한국정보처리학회 2024 정보처리학회논문지. 소프트웨어 및 데이터 공학 Vol.13 No.3
자연어 처리 분야에서 반어 및 비꼼 탐지의 중요성이 커지고 있음에도 불구하고, 한국어에 관한 연구는 다른 언어들에 비해 상대적으로 많이부족한 편이다. 본 연구는 한국어 텍스트에서의 반어 탐지를 위해 다양한 모델을 실험하는 것을 목적으로 한다. 본 연구는 BERT기반 모델인 KoBERT와 ChatGPT를 사용하여 반어 탐지 실험을 수행하였다. KoBERT의 경우, 감성 데이터를 추가 학습하는 두 가지 방법(전이 학습, 멀티태스크 학습)을적용하였다. 또한 ChatGPT의 경우, Few-Shot Learning기법을 적용하여 프롬프트에 입력되는 예시 문장의 개수를 증가시켜 실험하였다. 실험을수행한 결과, 감성 데이터를 추가학습한 전이 학습 모델과 멀티태스크 학습 모델이 감성 데이터를 추가 학습하지 않은 기본 모델보다 우수한 성능을보였다. 한편, ChatGPT는 KoBERT에 비해 현저히 낮은 성능을 나타내었으며, 입력 예시 문장의 개수를 증가시켜도 뚜렷한 성능 향상이 이루어지지않았다. 종합적으로, 본 연구는 KoBERT를 기반으로 한 모델이 ChatGPT보다 반어 탐지에 더 적합하다는 결론을 도출했으며, 감성 데이터의 추가학습이 반어 탐지 성능 향상에 기여할 수 있는 가능성을 제시하였다. Despite the increasing importance of irony and sarcasm detection in the field of natural language processing, research on the Koreanlanguage is relatively scarce compared to other languages. This study aims to experiment with various models for irony detection inKorean text. The study conducted irony detection experiments using KoBERT, a BERT-based model, and ChatGPT. For KoBERT, twomethods of additional training on sentiment data were applied (Transfer Learning and MultiTask Learning). Additionally, for ChatGPT,the Few-Shot Learning technique was applied by increasing the number of example sentences entered as prompts. The results of theexperiments showed that the Transfer Learning and MultiTask Learning models, which were trained with additional sentiment data,outperformed the baseline model without additional sentiment data. On the other hand, ChatGPT exhibited significantly lower performancecompared to KoBERT, and increasing the number of example sentences did not lead to a noticeable improvement in performance. Inconclusion, this study suggests that a model based on KoBERT is more suitable for irony detection than ChatG
미디어 텍스트 분석 기반의 공급망 리스크 모니터링 시스템의 개발
최동엽,서용원 한국생산관리학회 2023 한국생산관리학회지 Vol.34 No.4
오늘날의 기업들은 무역 갈등의 심화, 전염병, 경기 침체, 전쟁 및 각종 자연재해 등으로 인한 공급망 리스크에 노출되어 있다. 이러한 배경에서 공급망 리스크와 관련된 정보를 수집하고 동향을 파악하는 것의 중요성이 증대되고 있으며, 뉴스 기사와 같이 실시간으로 발생하는 미디어 텍스트를 분석하는 것은 공급망 리스크와 관련된 최신 정보를 빠르게 수집하는 유용한 방법으로 대두되고 있다. 하지만, 공급망 리스크와 관련된 텍스트를 분석하는 연구는 초기 단계이며, 최근 활용도가 증가하는 딥러닝 기반의 자연어 처리 기법을 적용한 텍스트 분석은 미비한 상황이다. 이에 본 연구에서는 뉴스 기사 분석을 활용하여 공급망 리스크와 관련된 정보를 수집, 도출하는 인공지능 기반의 공급망 리스크 모니터링 시스템을 개발한다. 이를 위해 사전 학습 언어모델인 KoBERT에 기반해 공급망 리스크 관련 기사만을 수집하는 필터링 모델을 수립하고, 수집된 기사의 공급망 리스크 유형을 LDA 토픽 모델링 기반으로 식별하여 학습 데이터를 구축하였다. 이후, BOW(Bag of Words)와 KoBERT를 사용한 딥러닝 기반의 공급망 리스크 분류 모델을 개발하여 수집된 기사의 공급망 리스크 유형을 예측하였다. 분석 결과, KoBERT 기반의 공급망 리스크 관련 기사의 필터링 정확도가 92.2%의 높은 성능을 보이는 것으로 나타났으며, 리스크 유형 분류 모델에서도 KoBERT 기반의 공급망 리스크 유형 분류 모델이 BOW 기반 모델에 비해 높은 분류 성능을 보이는 것을 확인하였다. Recently, companies are exposed to various supply chain risks such as intensified trade conflicts, epidemics, economic and geopolitical uncertainties, and natural disasters. Thus there is increasing importance in monitoring information related to supply chain risks. Analyzing real-time media texts, such as news articles, can be utilized for monitoring up-to-date information supply chain risks. However, researches regarding analyzing supply chain risk related text are in early stages, and researches to apply modern AI techniques such as deep learning-based natural language processing to supply chain risk texts are scarce. This study aims to develop a supply chain risk monitoring system that monitors and extracts information related to supply chain risks by analyzing news articles. To collect supply chain risk related articles a filtering model based on KoBERT is developed, of which risk types are identified based on LDA topic modeling to be utilized as the train data. To predict news articles’ risk types, two deep learning- based risk classification models are developed using BOW(Bag of Words) and KoBERT. The results showed high accuracy of KoBERT based model in filtering supply chain risk-related articles, and in the classification of supply chain risk types also KoBERT based model showed better performance than BOW based model.
KoBERT와 KoGPT2 기반의 대형언어모델과 딥러닝을 통합한 리뷰 유용성 예측모형
김은미,남승진,김태이,홍태호 한국지능정보시스템학회 2024 지능정보연구 Vol.30 No.2
AI 기술이 산업 전반에서 광범위하게 적용되면서 텍스트 데이터 기반의 대형언어모델이 높은 관심을 받고 있다. 대형 언어모델은 번역, 챗봇, 콘텐츠 생성 등과 같은 자연어 처리 분야에서 활발히 연구되고 있으며, 이커머스 분야에서도 고객 데이터 분석을 위해 사용되고 있다. 제품 및 서비스에 대한 사용 경험을 기반으로 사용자가 직접 작성하는 온라인 리뷰는 고객분석을 위한 중요한 자료이며, 대형언어모델의 활용은 텍스트로 작성되어 있는 리뷰 데이터의 의미 파악을 보다 정확 하게 할 수 있도록 한다. 본 연구는 SVM, 1D-CNN, 2D-CNN, CNN-LSTM의 딥러닝 모델과 대형 언어 모델인 KoBERT 와 KoGPT2로 리뷰 유용성 예측모형을 구축하고, 이를 통합하여 텍스트 내의 복잡한 의미가 반영된 리뷰 유용성 예측모 형을 제안한다. 구글 지도의 리뷰 데이터를 활용하였으며, 딥러닝 기법에서는 CNN-LSTM의 예측 성과가 72.74%로 가장 우수한 것으로 나타났다. KoBERT와 KoGPT2의 대형언어모델은 73.22%와 75.74%로 기존의 머신러닝 기법의 예측 모델 보다 대형언어모델을 기반으로 한 예측 모형이 우수한 성능을 보였다. 본 연구에서 제안한 딥러닝 기법과 대형 언어 모델을 통합한 통합모형에서는 76.37%의 정확도로 예측성과를 향상시켰으며, 통합모형은 텍스트의 의미를 보다 정확하게 반영하고, 예측성과를 향상시키며 예측모형의 안정성을 높일 수 있다. As AI technology is widely applied across industries, text data-based large language models are gaining significant attention. Large language models are actively researched in natural language processing fields such as translation, chatbots, and content creation, and are also used for customer data analysis in e-commerce. Online reviews, which are written directly by customers based on their experiences with products and services, are crucial for customer analysis, and leveraging large language models can help better understand the meanings embedded in these text reviews. This study proposes a review helpfulness prediction model by integrating deep learning models such as SVM, 1D-CNN, 2D-CNN, CNN-LSTM, and large language models KoBERT and KoGPT2, thereby reflecting the complex semantics within the text. Experiment results indicate that the CNN-LSTM model showed the best prediction performance at 72.74%. The large language models KoBERT and KoGPT2 achieved 73.22% and 75.74%, respectively, showing that prediction models based on large language models performed better than traditional machine learning models. The integrated model, combining the deep learning techniques and large language models proposed in this study, improved prediction performance with an accuracy of 76.37%, indicating that the integrated model can more accurately reflect the text’s meaning, enhance prediction performance, and improve model stability.
KoBERT 기반 비속어 검출 모델 및 FAST API 서버 구현
김영민(Young-Min Kim),박승민(Seung-Min Park) 한국전자통신학회 2024 한국전자통신학회 논문지 Vol.19 No.6
본 논문에서는 한국어 BERT(KoBERT)를 전이 학습하여 비속어가 포함된 문장과 그렇지 않은 문장을 구별하는 모델을 구축하고, 이를 Python의 FAST API를 이용하여 웹 서비스 형태로 구현한 연구 결과를 제시한다. 데이터 셋은 다양한 온라인 커뮤니티와 소셜 미디어에서 수집한 문장을 활용하였으며, 전처리 과정을 거쳐 비속어 여부로 라벨링 하였다. KoBERT를 기반으로 한 분류 모델을 구축하고, 전이 학습 기법을 통해 높은 정확도의 비속어 검출 성능을 달성하였다. 또한, FAST API를 이용하여 클라이언트로부터 POST 요청을 받아 텍스트 데이터를 처리하고, 비속어 여부를 반환하는 웹 서비스를 구현하였다. 본 연구는 KoBERT를 활용한 비속어 검출의 가능성을 확인하고, 실용적인 웹 서비스 구현을 통해 실제 적용 가능성을 제시하였다. 향후 연구로는 더 다양한 데이터 셋을 활용한 모델 성능 개선과 실시간 비속어 필터링 시스템 구현을 목표로 한다. This paper presents a study in which a model is built to distinguish between sentences containing profanity and those that do not, by applying transfer learning to KoBERT (Korean BERT). The model is implemented as a web service using Python’s FAST API. The dataset consists of sentences collected from various online communities and social media platforms, and after a preprocessing stage, the sentences were labeled based on the presence of profanity. A classification model was built using KoBERT, and by utilizing transfer learning techniques, high accuracy in profanity detection was achieved. Additionally, a web service was implemented using FAST API, which processes text data received through POST requests from clients and returns whether profanity is present or not. This study confirms the potential of using KoBERT for profanity detection and demonstrates the feasibility of practical application through the implementation of a web service. Future research will aim to improve model performance by utilizing more diverse datasets and to implement a real-time profanity filtering system.
토픽모델링 기반 고객 불만 분류 및 서비스 개선 방법 연구
원종수,배장원 한국정보통신학회 2026 한국정보통신학회논문지 Vol.30 No.1
정보통신기업에서는 초고속 인터넷, IPTV, 유선전화 서비스의 품질 관리를 위해 고객 불만(Voice of Customer, VoC) 데이터 분석의 중요성이 지속적으로 증가하고 있다. 본 연구는 한국어 특화 언어모델인 KoBERT와 BERTopic을 결합하여 110,809건의 고객 불만 데이터를 분석하고, 이를 서비스 품질관리 체계에 연계하고자 하였다. 제안한 방법은 KoBERT를 활용해 상담 텍스트를 문장 단위 임베딩으로 변환한 후, BERTopic 기반 토픽모델링을 적용하여 고객 불만 유형을 자동으로 군집화하는 방식이다. 이를 통해 전체 데이터를 인터넷 품질, TV 품질, A/S 접수 후 정상, 유선전화 품질, 고객 자가 장비 및 환경 A/S의 5개 주요 서비스 품질 카테고리로 구조화하였다. 실험 결과, KoBERT 임베딩 모델은 정확도 0.9907, F1-score 0.9906을 기록하여 기존 한국어 언어모델 대비 우수한 성능을 보였다. 또한 BERTopic 모델링 결과 전체 문서의 99.7%가 토픽에 할당되어 높은 군집 품질이 확인되었다. 이러한 결과는 고객 불만 데이터를 기반으로 한 선제적 품질관리와 운영 효율성 향상을 위한 데이터 기반 의사결정 지원 가능성을 시사한다. In the information and communications industry, analyzing customer complaints (Voice of Customer, VoC) is essential for managing the quality of high-speed Internet, IPTV, and fixed-line telephone services. This study integrates the Korean Bidirectional Encoder Representations from Transformers (KoBERT) model with Bidirectional Encoder Representations-based Topic Modeling (BERTopic) to analyze 110,809 customer complaint records and support service quality management. The proposed approach extracts sentence-level embeddings using KoBERT and applies BERTopic to automatically cluster complaint types. The dataset is structured into five service quality categories: Internet quality, TV quality, normal after A/S request, fixed-line telephone quality, and customer-owned equipment and environment A/S. Experimental results show that the KoBERT embedding model achieved an accuracy of 0.9907 and an F1-score of 0.9906, outperforming existing Korean language models. In addition, BERTopic assigned 99.7% of documents to topics, indicating high clustering quality. These results demonstrate that VoC-based topic modeling enables proactive quality management, data-driven decision-making, and improved operational efficiency in telecommunications services.
딥러닝 기반 [KoBERT-MLP] 제조업 위험성평가 빈도·강도 예측 모델 연구
정무환,어원석,이준원,오유라 사단법인 한국안전문화학회 2026 안전문화연구 Vol.- No.53
This study proposes a deep learning-based text analysis method utilizing Korean natural language processing to address the subjectivity and inconsistency issues that arise during manufacturing risk assessments. Conventional risk assessment practices rely on qualitative evaluation methods dependent on individual workers' experience and judgment, which frequently leads to inconsistent results even for identical work conditions. To address this limitation, the present study constructed a KoBERT-MLP text classification model using 16,380 risk assessment records collected from manufacturing workplaces, predicting the frequency and severity of hazard occurrence in a multi-class classification framework. The dataset was divided into training(80%) and test(20%) sets, and model performance was evaluated using Accuracy and Macro F1-score. To objectively position the proposed model, a TF-IDF-based Logistic Regression model was additionally employed as a baseline for comparative analysis. The analysis results indicated that severity prediction achieved slightly higher accuracy than frequency prediction. In terms of Accuracy, KoBERT-MLP outperformed the baseline model by 0.026 for frequency prediction and 0.052 for severity prediction, demonstrating the effectiveness of contextual embeddings in capturing the semantic content of risk assessment texts. However, Macro F1-scores were relatively low, which is attributed to the combined effect of uneven data distribution across grade levels and the model's tendency to converge toward majority classes. Confusion matrix analysis revealed that prediction errors predominantly occurred between adjacent grade levels, with a consistent tendency toward downward misclassification of high-risk grades. Notably, the baseline model achieved an F1-score of 0.345 for the highest severity grade, whereas KoBERT-MLP recorded an F1-score of 0.000 for the same grade, indicating a structural limitation in predicting rare extreme grades. This study demonstrates that worker judgment patterns in risk assessment can be quantitatively analyzed using descriptive text records, and may serve as foundational research for enhancing the objectivity and consistency of future risk assessments. However, the proposed model is limited to serving as a supplementary tool for mid-range risk grades, and expert judgment must be applied in parallel for final decisions on high-risk grades. Furthermore, this approach enables a data-driven understanding of workers' hazard perception and judgment criteria. It is anticipated that the standardization and consistency of risk assessment outcomes will contribute to enhancing overall organizational safety awareness, though the scope of contribution is limited to improving the consistency and objectivity of risk assessment, as safety culture itself is shaped by multiple interacting factors including education, organizational structure, and leadership.