
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
Sentiment-based evaluation of emotional experience in UX : a case study on Duolingo user feedback
YUSUFJONOV DIYORBEK IBROKHIM UGLI 국민대학교 비즈니스IT전문대학원 2025 국내석사
This thesis investigates how emotional user experiences in mobile language-learning applications can be systematically evaluated using large-scale app-store reviews. Focusing on 9,742 Google Play reviews of Duolingo collected between August 2024 and April 2025, the study develops the Sentiment–Emotion–Experience UX (SEE-UX) framework, which integrates sentiment analysis, emotion detection, transformer-based modeling, and feature-level UX mapping. Text preprocessing was followed by a combined lexicon-based sentiment analysis (VADER and TextBlob), emotion extraction using the NRC Emotion Lexicon, and a transformer-based sentiment classifier (DistilBERT) to enhance contextual accuracy. The results reveal a predominantly positive emotional landscape characterized by trust, anticipation, and joy, alongside concentrated negative emotions, particularly anger, sadness, and fear, around specific UX components such as the hearts system, ad interruptions, and subscription prompts. Transformer-based sentiment scores demonstrated stronger alignment with star ratings than lexicon methods, highlighting their value for capturing subtle dissatisfaction masked by polite or mixed-tone language. The study discusses how the SEE-UX framework could be extended to other language-learning applications and outlines cross-app validation as a promising direction for future work. The findings underscore the importance of emotion-aware UX design, offering practical insights for reducing friction, improving motivational features, and designing more empathetic monetization strategies. Overall, the study demonstrates that emotional analytics, when combined with UX theory and modern NLP methods, provide a powerful approach for evaluating and improving digital learning experiences. 본 연구는 모바일 언어 학습 애플리케이션에서의 사용자 감정 경험을 앱스토어 리 뷰를 기반으로 체계적으로 분석하고 UX 개선으로 연결하는 방법을 탐구한다. 사례 연구 대 상으로 Duolingo의 Google Play 영어 리뷰를 수집하여 전처리한 뒤, VADER와 NRC 감정 어휘 기반 분석과 함께 DistilBERT 기반 문맥적 감성 모델을 적용하였다. 또한 리뷰 내 기 능 관련 표현을 활용해 감정 신호를 구체적 UX 구성요소(예: 하트, 광고, 구독, 스트릭)와 연결하는 SEE-UX(Sentiment–Emotion–Experience for UX) 프레임워크를 제안하였다. 분 석 결과, 전반적으로 긍정 정서가 우세했으나 하트(페널티), 광고 노출, 구독 유도와 같은 기 능은 분노·슬픔·두려움 등 부정 정서를 유발하는 주요 요인으로 나타났다. 또한 변환기 기반 모델은 어휘 기반 접근보다 미묘한 불만과 맥락을 더 잘 포착하여 별점과의 정렬성이 높았 다. 본 연구는 대규모 사용자 피드백에서 감정 기반 UX 인사이트를 도출하고, 제품팀이 실 행 가능한 개선 지침으로 전환할 수 있는 절차적 틀을 제공한다.
Sentiment Analysis for Government Services in Indonesia: A Systematic Literature Review Noval Hudiya
As one of the tasks in Natural Language Processing (NLP), sentiment analysis enables computational understanding of sentiment from textual public opinion data. With the rapid growth of public opinion data available online, sentiment analysis offers a cost-effective and potentially real-time alternative for evaluating government services. Although research on sentiment analysis has increased significantly in Indonesia in recent years, it still faces major challenges related to linguistic diversity, limited datasets, and low representation in global NLP datasets. While the interest in this field is growing, and the strategic alignment with Indonesia’s National Artificial Intelligence Strategy 2020-2045, to our knowledge, no review study has yet examined how sentiment analysis has been applied in Indonesia’s government services domain. This study fills that gap by performing a systematic literature review of studies published between 2021 and 2025, following the PRISMA framework guidelines across Scopus, ScienceDirect, and IEEE Xplore databases. A total of 107 studies met the inclusion criteria and were reviewed deeper to synthesize methodological trends, dataset sources, and domain applications of sentiment analysis in this field. The review finds machine learning approaches such as Naïve Bayes and SVM remain dominantly used in current practice, while deep learning models like LSTM and BERT start to grow. Twitter is often used as a primary data source, and the health domain, especially COVID-19-related studies, is the largest distribution of studies. Key challenges were found, including data imbalance and representativeness limitations, limited Indonesian government-specific linguistic resources, the presence of bot-generated data, a lack of service integration, and the absence of studies that measure sentiment analysis impact in real settings. This study contributes by identifying methodological and contextual gaps that can inform researchers in future sentiment analysis research in Indonesia’s government service sector. 자연어처리(NLP)의 핵심 과제 중 하나인 감성 분석은 텍스트로 표현된 대중 의견 데이터로부터 감정을 계산적으로 이해할 수 있도록 한다. 온라인에서 공개되는 의견 데이터가 빠르게 증가함에 따라, 감성 분석은 정부 서비스를 평가하는 데 비용 효율적이면서도 실시간에 가까운 대안을 제공한다. 인도네시아에서 감성 분석 연구는 최근 크게 증가했지만, 언어적 다양성, 제한된 데이터셋, 글로벌 NLP 데이터셋에서의 낮은 대표성 등 여러 주요 도전에 여전히 직면해 있다. 이러한 분야에 대한 관심이 높아지고 있으며 인도네시아의 「국가 인공지능 전략 2020–2045」와도 전략적으로 맞닿아 있음에도 불구하고, 현재까지 인도네시아 정부 서비스 분야에서 감성 분석이 어떻게 적용되어 왔는지를 검토한 리뷰 연구는 존재하지 않는 것으로 보인다. 본 연구는 Scopus, ScienceDirect, IEEEXplore 데이터베이스를 대상으로 PRISMA 가이드라인을 준수하여 2021년부터 2025년까지 발표된 연구들을 체계적으로 검토함으로써 이러한 공백을 메운다. 총 107편의 연구가 선정 기준에 부합했으며, 이들 연구를 심층 분석하여 감성 분석의 방법론적 경향, 데이터셋 특성, 그리고 분야별 적용 양상을 종합적으로 정리하였다. 분석 결과, 나이브 베이즈와 SVM과 같은 전통적인 머신러닝 기법이 여전히 널리 사용되고 있으며, LSTM과 BERT와 같은 딥러닝 모델의 활용도 점차 증가하고 있는 것으로 나타났다. 데이터 출처로는 트위터가 가장 많이 사용되고, 보건 분야, 특히 COVID19 관련 연, 가 가장 큰 비중을 차지하였다. 또한 데이터 불균형 및 대표성 부족, 인도네시아 정부 서비스에 특화된 언어 자원의 부재, 봇 생성 데이터의 존재, 서비스 간 통합 부족, 감성 분석이 실제 공공 서비스 제공에 미치는 영향을 평가한 연구의 부재 등 여러 중요한 한계점도 확인되었다. 본 연구는 이러한 방법론적·맥락적 연구 공백을 제시함으로써 향후 인도네시아 정부 서비스 분야에서의 감성 분석 연구를 발전시키는 데 필요한 방향성을 제공한다.
기계학습을 이용한 Aspect-Based Sentiment Analysis 기반 전기차 요소별 사용자 감성 분석 및 예측 모델링
In this study, we extract main components and attributes, which are the main aspects of Electric Vehicle by analyzing User Experience based on Aspect-Based Sentiment Analysis (ABSA) using machine learning, overcoming the problems accompanying in this process such as Data Imbalance and insufficient user reviews by making use of non-label data. In addition, we find the contributing factors affecting user’s sentiments by figuring out the relationship between user's sentiment to each aspect extracted and detailed specifications of Electric Vehicle with regression. Based on the ABSA method, and we perform data collection, data preprocessing, feature engineering, Aspect Extraction, modeling for sentiment analysis, and evaluating user sentiment to each aspect in sequence. For data collection, a total of 5,065 label data, which is evaluated with a 5-point scale by users, was collected from representative car forums. At the same time, in order to overcome the shortage of data and data imbalance, approximately 210,000 items of non-label data are collected from Youtube.com, of which 6,488 items were selected by filtering with limited to the user experience related only. And then, feature engineering is performed with effective embedding methods of distributed representation after data pre-processing. The analysis phase is mainly divided into two processes: Aspect Extraction and Sentiment Analysis. First of all, TextRank and Naïve methods were used as an unsupervised method and an extractive approach for Aspect Extraction. Then, in order to implement a sentiment classification model based on supervised learning with high performance, we built a machine learning model that trains the truncated text composed of one or two sentences at the beginning of a review text with a label and make it improved by means of semi-supervised learning. With the model trained, we are able to perform aspect-wise sentiment analysis by conducting sentiment analysis on the sentence that including the selected aspect term. Further, we find detailed specifications of vehicle that have an influence on user sentiment as contributing factors that affects user’s sentiment. As a result, 16 categories of main aspects were extracted, eight key EV Components & eight key Human Factor Attributes, of which the users are likely to be positive to “Acceleration, Room, Interior, Power, Safety, Ergonomics, Price, Power” and negative to “Seat, Battery, Charge, Noise, Winter, Ice”. In sentiment analysis, the CNN model showed the highest performance in sentiment classification. Therefore, through semi-supervised learning using CNN, label propagation was performed among non-label data, giving the pseudo label to only the data with a high classification probability more that 80%, resulting in improvement in performance of the CNN model. Lastly, we confirmed the high classification accuracy of the deep learning model for predicting the user’s sentiment of the sentences. In addition, with regard to aspect-wise sentiment analysis, there was a tendency to predict the user’s sentiment similarly between machine learning based and lexicon-based, which showed machine learning based model is robust as much as lexicon-based. In conclusion, it was shown that more diverse topics and unbiased opinions could be extracted through aspect-wise analysis than review-wise. In addition, we verified that the imbalance problem could be overcome by over-sampling Finally, a more effective UX analysis framework for the products that have not sufficient user reviews was proposed by taking advantage of non-label data with semi-supervised learning. 본 논문에서는 전기차를 대상으로 기계학습을 이용한 Aspect-Based Sentiment Analysis(ABSA) 기반 사용자 리뷰 분석을 통해, 차량의 주요 요소(Aspect)인 부품(Components) 및 특징(Attributes)을 추출하고, 추출된 각 요소에 대한 사용자 감성 예측 모델링 기반의 UX 분석 프레임워크(Framework)를 구현하여 기존의 인터뷰 및 설문조사와 유사한 수준의 사용자 의견을 얻는 것을 주요 목표로 한다. 이 과정에서 수반되는 데이터 불균형(Data Imbalance) 문제를 오버샘플링(Oversampling)을 통해 극복하고, 사용자 리뷰 부족 문제 극복을 위해 레이블이 없는(Non-label) 데이터를 활용하는 방법을 제안한다. 더불어 추출된 Aspect에 대한 차량 세부 스펙과 사용자 감성 간의 관계성 확인을 통해 감성에 영향을 주는 요소(Contributing Factor)를 찾는다. 연구 방법은 ABSA의 큰 틀을 활용하며, 크게 데이터 수집, 전처리 및 Feature 생성, 요소 추출(Aspect Extraction) 및 감성 분석(Sentiment Analysis)을 위한 모델링, 그리고 요소 별 사용자 감성 분석 순서로 진행하였다. 데이터 수집은 대표적인 자동차 포럼에서 사용자 만족도가 5점 척도로 평가된 Label 데이터 총 5,065개를 수집하였고, 데이터 부족 문제를 극복하고자 Youtube.com에서 Non-label 데이터를 약 21만개 수집하였으며 이 중 User Experience 관련 어휘가 포함된 리뷰로 한정하여 총 6,488개를 선별하였다. 이후 수집 데이터의 전처리 및 분산 표현(Distributed Representation)을 통한 효과적인 임베딩 과정을 거쳐 특징(Feature)을 생성하였다. 분석은 크게 두 가지 줄기로써, 요소 추출(Aspect Extraction)과 감성 분석(Sentiment Analysis)로 나뉜다. 요소 탐지를 위해 비지도적 방법(Unsupervised Method)이자 추출적 접근 방법(Extractive Approach)으로써, TextRank와 Naïve Method를 활용하였다. 그 다음 지도학습(Supervised Learning) 기반의 문장 감성 분류 모델을 구현하고자, Label이 있는 리뷰 텍스트 서두의 한 두문장으로 구성된 절단된 텍스트를 학습시킨 모델을 구축하였고, 준지도학습을 통해 더 나은 성능의 모델을 구현하고자 하였다. 이를 바탕으로 선정된 Aspect가 포함된 문장에 대한 감성 분석을 실시함으로써 요소별 감성 분석을 진행하고, 더불어 사용자 감성에 영향력 있는 차량 세부 스펙을 찾아 Contributing Factor를 발굴하고자 하였다. 연구 결과로써, 요소 추출(Aspect Extraction)로는 총 16개 카테고리의 주요 Aspects(8개의 주요 전기차 구성 요소와 8개의 주요 Human Factor 특성)가 추출되었는데, 이 중 사용자는 “Acceleration / Room / Interior / Power / Safety / Ergonomics / Price / Power”에 대해 긍정적이며, “Seat / Battery / Charge / Noise / Winter / Ice”에 대해 다소 부정적임을 확인하였다. 감성 분석(Sentiment Analysis)에서는 CNN 모델이 리뷰 단위 감성 분류에 있어 가장 높은 성능을 보였다. 따라서 CNN을 활용한 준지도학습(Semi-Supervised Learning)을 통해 Non-Label Data 중 80% 이상의 분류 확률이 높은 데이터 위주로 Pseudo Label을 부여하였고, 이를 포함한 전체 데이터를 재학습을 거치는 방법으로 모델의 성능 향상을 확인하였다. 또한 추출된 요소가 포함된 문장 단위 감성 분류에 대하여, 기계학습 모델 기반으로 결과와 Lexicon 기반 감성 분류 결과 간 17개 Aspect 중 14개가 예측 방향성이 일치함을 확인함으로써, 기계학습 기반 감성 분류 모델의 타당성을 간접적으로 확인하였다. 마지막으로 샘플 검증을 통해 본 연구에서 학습된 딥러닝 모델의 높은 분류 정확도를 확인하였는데, 딥러닝 모델이 단어 의미 이상으로 문장 문맥을 파악하여 긍정/부정 분류하였음을 확인하였다. 결론적으로 Aspect 기반의 문장단위 분석을 통해 보다 더 다양한 토픽과 편향되지 않은 의견을 추출할 수 있음을 보였다. 더불어 리뷰 데이터를 Over-sampling을 하여 Data Imbalance 문제를 접근함으로써 온라인 리뷰의 긍정 편향성을 극복하고, Semi-Supervised Learning을 통한 Non-Label Data 활용 방법을 통해 사용자 평가가 많이 부족한 제품에 대해 보다 효과적인 UX 분석 프레임워크를 제안하였다.
윤현애 연세대학교 일반대학원 2025 국내박사
This study investigates how appraisals of target objects are linguistically realized in Korean product reviews under the goal of sentiment analysis. From the Functionalist linguistics perspective, language exists as a tool for communication, and speakers may choose particular lexis, grammatical structures, or broader patterns to convey their intentions. In other words, meaning emerges from “language in use,” placing greater emphasis on semantic and functional aspects rather than formal ones. This study begins with the question: which expressions do speakers employ to reveal their own cognitive attitudes toward a target? Speakers may directly encode appraisals or judgments—using descriptors such as “good,” “bad,” or “beautiful”—but they also often utilize fact-stative expressions (e.g., “there are many flowers”) to convey an evaluation of a target’s aesthetic state indirectly. Notably, when such fact-stative expressions occur in contexts that communicate experiences of tourist sites, they can be interpreted as appraisals. This observation suggests that whether a factual description is construed as an appraisal depends on shared knowledge and contextual prerequisites among discourse participants. Indeed, delineating precisely which linguistic expressions in discourse qualify as evaluations of a target is a highly demanding undertaking. J. Martin’s systemic-functional Appraisal Theory further exemplifies this challenge, as it continually uncovers the intricate and arduous nature of mapping the correlation between linguistic expressions and evaluative functions. In this way, the correlation between a given linguistic expression and its appraisal in discourse must take into account a variety of conditions—such as the type of appraisal target, the language community’s expectations of that target, the contextual factors surrounding the appraisal, and the relationships among interlocutors. However, because these factors are highly variable and complex, systematically defining and analyzing them is exceedingly difficult. However, under the goal of sentiment analysis, it is feasible to identify the linguistic characteristics of speakers’ direct and indirect appraisals of a target. Because sentiment analysis focuses on publics’ positive and negative evaluations of items, the scope of appraisal targets can be narrowed to those entities or issues that attract public attention. In particular, when the discourse under investigation consists of product reviews, shared knowledge among interlocutors is confined to a specific product, and any omitted sentence elements can often be recovered easily from the product-review context. Moreover, since sentiment analysis centers on “positive–negative evaluation,” the difficulty of determining which expressions qualify as appraisals is somewhat alleviated. In other words, sentiment-analysis research conducted within the contextual constraints of product-review discourse—where both the appraisal target and context are relatively circumscribed—is well suited to exploring the linguistic realization of appraisal in discourse. Meanwhile, much of the existing work in sentiment analysis has concentrated on individual words, constructing lexicons of positive and negative terms. In many dataset-building procedures, the text is first morphologically analyzed and POS-tagged, after which verbs and adverbs are extracted and assigned polarity labels. However, in actual language use, speakers do not restrict their appraisals to single words—word-focused methods centered on verbs and adverbs therefore have inherent limitations. Moreover, such lexicon-centric approaches often fail to account for instances where a potentially positive term’s evaluative force is neutralized or altered by specific grammatical constructions. Consequently, it is essential to move beyond word-level techniques and to analyze both lexical and grammatical expressions that realize appraisal in discourse, investigating their full range of semantic characteristics. However widespread generative language models (LLMs) have become, the need for meticulously constructed sentiment‐analysis datasets remains high when accuracy and practical applicability are taken into account. Because appraisals are inherently context-dependent—even when defined simply as positive or negative evaluations—relying solely on general-domain LLMs limits one’s ability to perform fine-grained, domain- and context-specific sentiment analysis. Accordingly, this study presents a comprehensive inventory of lexical and grammatical expressions that convey positive and negative sentiment, together with an analysis of their semantic and functional characteristics. The resulting catalog can serve as a set of seed words and constructions for researchers and practitioners across both academic and industrial contexts. More broadly, our findings offer a valuable starting point for understanding the linguistic features through which speakers express appraisal in discourse. This study defines appraisal as the act of judging some aspect of an existing target as positive or negative. We broaden our scope beyond predicates and adverbs to include all parts of speech—including nouns—and extend our analysis to phrases and clauses. The objectives of the present study are twofold: (1) To identify which semantic properties of lexical items in product‐review discourse signal appraisal. (2) To determine how grammatical meanings in product‐review discourse express or modulate appraisal. In Chapter 2, we first survey both domestic and international research on sentiment and emotion analysis, alongside lexical classification studies conducted in Korean linguistics and language informatics. This review reveals that investigations into sentiment and emotion analysis have been pursued continuously across multiple arenas—from individual and corporate projects to national initiatives. However, dataset construction for sentiment analysis remains largely domain-specific. Accordingly, there is a pressing need to conduct cross-domain analyses that distinguish between (a) expressions that consistently convey positive or negative evaluations regardless of domain and (b) expressions whose evaluative force is confined to particular domains. Although the importance of such a distinction has been emphasized repeatedly, few studies have empirically attempted to extract truly domain-independent evaluative expressions through multi-domain analysis. Furthermore, certain classification categories—such as “price” or “design”—may function as universal sentiment dimensions across domains, underscoring the necessity of rigorous, empirical research in this area. Secondly, it is somewhat regrettable that many sentiment‐analysis studies stop at the stage of listing lexical items without providing semantic justification for why those expressions convey positive or negative appraisal. If we understand the underlying semantic motivations, we can more effectively extend the inventory to include semantically similar expressions, thereby constructing sentiment lexicons with greater efficiency. Moreover, by examining how grammatical constructions themselves encode or modulate evaluative meaning, we can further enhance the accuracy of sentiment‐analysis systems. Thirdly, although recent research has advanced toward fine‐grained, attribute‐based sentiment analysis, many studies still do not thoroughly address the problem of attribute extraction. As Kim Hansaem(2022) observes, distinguishing between entities and their evaluative attributes in annotated corpora is a challenging undertaking that demands dedicated investigation. By tackling this distinction head‐on, we can better support the development of high‐precision sentiment‐analysis resources. Drawing on Korean lexical‐classification studies, we identified several shared criteria for distinguishing among parts of speech. In verb‐classification research—whether focused on case‐frame patterns or semantic properties—scholars commonly differentiate between eventivity and stativeity, as well as the animacy of the subject and whether the verb describes the subject’s psychological experience. In noun‐classification work, the prevailing distinctions concern whether a noun denotes a concrete entity (human or object) versus an abstract concept, and whether it expresses relational meaning with other lexical items. On this basis, our study adopts [eventivity], [stativeity], and the presence or absence of psychological‐state description as key semantic features for categorizing lexical items. Furthermore, in line with ontological and thesaurus‐based knowledge‐base practices, we abstract away from traditional part-of-speech boundaries, classifying expressions purely by their conceptual and semantic properties. We then examine how these semantically motivated categories relate to the realization of appraisal in product‐review discourse. In Chapter 3, this study describes the research methodology. We constructed an 82,289-words, multi-domain corpus for sentiment analysis by extending the National Institute of Korean Language’s Attribute-Based Sentiment Analysis Corpus 2021 with additional product-review data. Four domains—beauty products (small-sized goods), home appliances (large-sized goods), lodging establishments (place-based services), and films (content products)—were each normalized to approximately 20,000 words to ensure balanced representation. To maximize the density of appraisal expressions, non-evaluative sentences from the primarily blog-style NIKL corpus were filtered out. In the beauty-product and film domains, roughly 70 percent of the original NIKL data were retained and supplemented with randomly sampled Naver beauty-product reviews from 2023 and Naver movie reviews from January through June 2024. For the home-appliances domain, we randomly sampled Naver product reviews from 2019, and for the lodging-establishment domain, we collected randomly sampled TripAdvisor reviews from 2023. This rigorously curated, multi-domain corpus enables a systematic exploration of appraisal expressions across varied product contexts while preserving both comparability and domain-specific richness. I define the core terminology of this study. I distinguish sentiment expressions from appraisal expressions as follows. Sentiment expressions are those linguistic items that, by themselves, overtly convey a positive or negative evaluation of a target. In contrast, appraisal expressions form a broader category: they comprise any lexical or grammatical expression that, when accompanied by specific morphological or syntactic markers (e.g., attribute modifiers, evaluative particles), can function as a sentiment expression. Put differently, because our framework is grounded in attribute-based sentiment analysis, we treat sentiment expressions as the concatenation of an attribute expression and an appraisal expression—only in their combination does a full evaluative meaning emerge. Also, this study also delineates the distinction between entities and attributes. Although entities ordinarily refer to a product’s components and attributes to its properties, this binary proves difficult to uphold in practice. In attribute-based sentiment analysis, attributes take precedence: they are the specific aspects or elements of a target that speakers judge positively or negatively. Consequently, the scope of an attribute varies with the analyst’s evaluative objectives and is formalized via the sentiment-analysis taxonomy. For the present study, we therefore define attributes as the concrete targets or elements that bear positive or negative evaluative polarity within attribute-based sentiment analysis. I outline a seven-stage analysis procedure: (1) Preprocessing: segmented the corpus into sentences—each assigned a unique index—and imported them into Excel to align with my clause-level analytical unit. (2) Framework Development: established an attribute-based sentiment-analysis schema, defining the criteria for identifying appraisal expressions and tagging attributes. (3) Polarity Annotation: applied this schema to assign positive or negative polarity to each evaluative instance. (4) Expression & Attribute Tagging (5) Distinguished into Attribute-Implicit and Attribute-Explicit expressions I introduced two novel appraisal categories: - Attribute-Implicit Expressions (속성내포형): Surface forms that do not overtly mark the attribute but whose semantics imply it (e.g., moisturizing). - Attribute-Explicit Expressions (속성명시형): Forms that lexically specify the attribute being evaluated (e.g., rich in hydration, good moisturizing effect). Although steps 2–4 were conducted interactively to ensure alignment between my schema and the data, this structured pipeline enabled a rigorous, reproducible analysis of appraisal expressions in product-review discourse. I term linguistic expressions denoting attributes—such as 수분감 (moisture sensation) and 보습력 (moisturizing power)—as attribute expressions. To qualify as an attribute expression, I require three conditions to be met: (1) It must constitute a separable linguistic unit at least at the phrase level. (2) It must represent a specific positive or negative aspect of the product. (3) The expression alone must not suffice to determine a positive, negative, or neutral evaluation. Next, I conducted a correlation analysis between the semantic properties of lexical and grammatical expressions and their evaluative force. The central question was: “On what basis can a sentiment expression be interpreted as conveying positive or negative appraisal?” For example, the positive appraisal of expressions like “상품 좋다” (“the product is good”) and “디자인 예쁘다” (“the design is pretty”) rests on the fact that 좋다 and 예쁘다 carry positive lexical meanings. In contrast, the positive appraisal of “사용할 만하다” is licensed by the grammatical construction ‘-(으)ㄹ 만하다’ (pronounced [-(eu)l manhada]), which grammatically encodes the positive sense of “to be worth doing.” Thus, an expression’s evaluative polarity in discourse may derive either from its inherent lexical semantics (as in the first case) or from the semantic contribution of a grammatical marker (as in the second). In this study, I analyzed whether each extracted sentiment expression’s evaluative basis is attributable to its lexical-semantic properties or to the semantics of its grammatical construction. When the basis for sentiment expressions resides in lexical items, I first conducted semantic‐feature grouping, followed by an analysis of the correlation between semantic features and appraisal. Semantic‐feature grouping began with initial lexical clusters derived from existing semantic‐attribute classification studies; I then iteratively refined these clusters by typologically analyzing the list of sentiment expressions extracted from the corpus. In Chapter 5, I present the correlation‐analysis results by dividing the clusters into three types. By contrast, for the correlation analysis between grammatical expressions and appraisal realization, I did not perform semantic‐feature grouping. Instead, I directly analyzed the correlation across all grammatical constructions observed in the data. This approach reflects the tendency—already noted in this study—for lexical items, rather than grammatical forms, to bear the primary evaluative load. For the grammatical analysis, I relied on the Kim et al (2005) dictionary of grammatical constructions. In Chapter 4, I present the results of sentiment analysis across the four selected domains. This chapter details the factors considered in constructing the sentiment classification schema, the attribute analysis, and the domain‐specific inventory of attribute‐implicit and attribute‐explicit sentiment expressions. Its purpose is not merely to report the evaluative profiles of beauty products, home appliances, lodging establishments, and films, but to reveal how reviewers perceive the contextual framing of each product type and linguistically realize their appraisals. Overall, the material‐goods domains (beauty products and home appliances) and the place‐service domain (lodging establishments) share common evaluative dimensions—such as price, design (aesthetic qualities), tactile experience, and service quality. In contrast, the film domain, as a content product, exhibits minimal overlap with these categories, aside from “fame” and “target audience.” Moreover, the film reviews demonstrate that a complete understanding of cinematic evaluation often requires integrating emotion analysis alongside sentiment analysis. This chapter’s findings offer practical guidance for analysts on the domain‐specific considerations essential to accurately capturing appraisal expressions in product‐review discourse. In Chapter 4, I move beyond a mere report of sentiment‐analysis results for the four domains. Instead, I introduce and justify the novel distinction between attribute-implicit and attribute-explicit expressions—thereby addressing long-neglected questions of attribute definition, delimitation, and extraction. As attribute-based sentiment analysis gains traction, the critical issue becomes how analysts should conceptualize attributes, construct classification schemas, and handle product-review corpora in practice. - Attribute-implicit expressions can be directly mapped to domain-specific appraisal categories within each sentiment‐classification framework. - Attribute-explicit expressions, as composites of an attribute expression and an appraisal expression, suggest a two-stage lexicon approach: (1) build separate attribute and appraisal dictionaries, then (2) establish domain-appropriate mapping rules between them to efficiently manage and extend both resources. Moreover, by applying synonym-expansion techniques to both dictionaries, one can enrich coverage with expressions not attested in the original corpus. These strategies collectively provide a scalable, systematic method for capturing evaluative language in diverse product‐review contexts. In Chapter 5, I investigate how the semantic characteristics of both lexical and grammatical expressions correlate with positive and negative appraisal in product‐review discourse. First, I identify two classes of lexical items that consistently function as evaluative markers: affective expressions (e.g., terms denoting pleasure or displeasure) and sensory expressions (e.g., descriptors of texture or scent). Affective expressions further subdivide into domain-independent items that convey approval or disapproval regardless of context and domain-dependent items whose evaluative force varies by product category. Sensory expressions—since they reflect the reviewer’s direct perceptual experience—almost invariably participate in appraisal across all four domains. Next, I turn to lexical expressions whose evaluative status depends on contextual or syntactic conditions. Appearance descriptors such as aesthetic or cleanliness terms generally carry appraisal, whereas descriptions of shape or passive constructions only do so when paired with specific attribute expressions or within particular domains. Property-descriptive terms tend to map onto concrete product attributes and exemplify the attribute-implicit category introduced earlier. Emotion terms likewise display varying behavior: core feel-good words (joy, fun, awe, relief, confidence) and core feel-bad words (resentment, disgust, aversion, embarrassment) appear universally as evaluative, while a second set of emotions (trust, hope, gratitude, regret, worry) requires morpho-syntactic support to function as appraisal. In the film domain, additional emotion words (empathy, anger, sadness, fear, surprise, bittersweetness) bridge sentiment and emotion analysis. Perceptual-cognitive expressions—aside from a few like “understand”—also need co-occurrence with attribute markers or contextual cues to convey evaluation, and eventivity expressions that imply repeat purchase or ongoing use consistently signal positive appraisal. Finally, I examine non-lexical items—comparatives, degree modifiers, and material descriptors—that themselves lack inherent polarity but serve as functional operators completing the appraisal when combined with attribute expressions. Together, these findings clarify the precise conditions under which various linguistic features realize positive and negative evaluations in product‐review discourse, offering a nuanced account of how speakers linguistically encode appraisal across diverse domains. In Chapter 6, I investigate the correlation between grammatical constructions and evaluative polarity in product‐review discourse. I first identify a set of grammatical markers that consistently convey positive or negative appraisal: ‘-(으)ㄹ 만하다’, ‘-어/아 보세요’, ‘-(으)ㄹ 수 있다’, ‘-(으)면 되다’, ‘-어/아도 되다’, and ‘-기는 하다’. Each of these constructions either inherently encodes a positive meaning or is employed in evaluative contexts so reliably that it functions as a grammatical appraisal expression. Beyond these unambiguous cases, I show that certain mood and modality markers—those expressing unmet conditions ([condition]), desire ([wish]), or volition ([will])—serve evaluative functions only when combined with particular lexical items. For example, ‘-(으)면’ in the conditional yields a positive appraisal when its protasis is negated and its apodosis contains a regret‐laden term; when followed by simple approval verbs like ‘좋다’ or ‘괜찮다’, it signals weak positive or neutral appraisal. The necessity marker ‘-어/아야’, when paired with positively valenced descriptions, conveys neutral-to-weakly negative evaluation by implying obligation. The wish construction ‘-(으)면 좋겠다’, combined with positive appraisal verbs, paradoxically delivers a negative evaluation by highlighting the absence of the desired state; its variant used in film‐review contexts (e.g., “I wish there were a sequel”) functions as a positive appraisal. Volitional forms (e.g., ‘-겠-’, ‘-어/아야겠다’, ‘-(을) 것이다’, ‘-(으)ㄹ게요’) become evaluative only when they express intention to repurchase or continue use. I also explore how case particles contribute to evaluation. The comparative particle ‘만큼’, when used with expectation nouns or price attributes, carries evaluative weight, while the limiting particle ‘만’, when followed by action‐oriented verbs or positive appraisal adjectives, distributes positive appraisal to the specified attribute and negative appraisal elsewhere. The copular particle ‘이다’ assists in realizing appraisal by asserting the presence of a valued attribute (e.g., “has a sea view”). Lastly, I document several context‐dependent appraisal constructions—such as the imperative ‘-(으)세요’, the retrospective ‘-(으)ㄹ 텐데’ and ‘-(으)ㄹ걸’, the obligation marker ‘-어/아야 하다’, and aspectual ‘-어/아 버리다’—noting that some (e.g., causatives ‘-게 하다’, ‘-게 만들다’) are evaluative only in the film domain or when combined with perceptual‐cognitive terms, and that certain markers (e.g., ‘-(으)ㄹ 때’) can neutralize a preceding appraisal. Through this comprehensive analysis, Chapter 6 delineates the precise grammatical conditions under which Korean evaluative meanings emerge in product‐review discourse. In Chapters 5 and 6, I examine how distinct semantic features shape positive and negative polarity judgments from the perspectives of lexis and grammar, respectively. In Chapter 5, I demonstrate that simple affective terms and sensory descriptors invariably signal evaluative polarity, making them prime candidates for inclusion in appraisal lexicons. Aesthetic- and cleanliness-related appearance descriptors, together with certain property descriptors and emotion terms whose polarity is unambiguous, likewise function consistently as evaluative markers. By contrast, the remaining appearance and property descriptors, as well as many emotion, perceptual-cognitive, and eventivity expressions, require specific contextual or attribute-expression conditions before their polarity can be resolved; this insight informs domain-sensitive sentiment‐analysis strategies. Notably, my finding that comparative, degree-modifier, and material expressions serve primarily as functional operators—while the attribute expressions themselves determine evaluative value—highlights the existence of different classes of attribute expressions: some merely denote product features or components, while others, by virtue of being mentioned, presuppose the presence of an attribute and thus influence polarity judgment. Distinguishing between these attribute classes will be an important focus for future research. Chapter 6 of this study is significant in that it ventures into a largely unexplored territory: the relationship between grammatical constructions and positive–negative appraisal judgments. While most sentiment-analysis research has focused on lexical items and produced extensive word lists, I demonstrate that certain purely grammatical markers—though devoid of inherent evaluative meaning—become integral components of sentiment expressions when combined with other elements. These constructions not only contribute to the assignment of polarity but, in some cases, actually trigger polarity shifts or neutralization. This finding underscores the need for future sentiment-analysis frameworks and datasets to extend beyond the lexicon and systematically incorporate grammatical phenomena. Nonetheless, a limitation of Chapter 6 is that it stops at cataloguing the grammatical appraisal expressions that influence polarity decisions. Although it is plausible that these constructions’ modal nuances play a crucial role in their evaluative function, the present study does not undertake a deep examination of their modality features. Subsequent research that probes these morphosyntactic subtleties would not only enrich our understanding of how speakers linguistically realize appraisal but also enhance the precision and sophistication of sentiment-analysis methodologies. When humans evaluate a target, they may employ direct lexical or expressive means—such as affective or emotion terms that reveal the speaker’s judgment or psychological state—but they also frequently use ostensibly factual descriptions (e.g., appearance or property descriptors) to convey positive or negative appraisal. Furthermore, under certain contextual conditions, expressions of obligation or volition can likewise function as evaluative markers. While lexical semantic features exert the primary influence on the realization of evaluative polarity, grammatical semantics also contribute to the expression or neutralization of such judgments; hence, both lexical and grammatical meanings must be considered in appraisal analysis. Many expressions—such as fact‐descriptive language, eventivity terms, and volitional or deontic constructions—are interpretable as evaluative only by virtue of interlocutors’ expectations and the shared knowledge of the discourse community. The sentiment analysis presented here ultimately concerns the study of how speakers’ evaluations are linguistically manifested, underscoring that any research engaging with the semantic dimension of Korean text must be grounded in rigorous linguistic inquiry. Because sentiment analysis interrogates both the latent meanings of expressions and the speaker’s intent, careful interpretation of surface‐level language forms is inseparable from linguistic theory. By integrating linguistics’ conceptual frameworks and empirical findings with the goals of sentiment analysis, this study provides a model for the definition, development, and application of analytical methods required to bridge theory and practice. 기능주의 관점에서 언어는 의사소통의 도구로서 존재하며 ’사용으로서의 언어(language in use)’일 때 언어의 실제 의미를 알 수 있다고 본다. 그렇다면 화자는 어떤 표현을 사용하여 대상에 대한 자신의 인지적 태도를 드러내는가? 본 연구는 이 문제를 탐구하기 위해 시작되었다. 사실상 담화에서 어느 언어 표현까지가 대상에 대한 평가를 나타내는가의 문제를 규명하는 일은 평가 대상 및 맥락 요소, 화청자 간의 공유 지식이 매우 다변적이고 복합적으로 얽혀 있다는 점에서 매우 어려운 과업이다. 그러나 상품평 담화 대상의 감성분석(sentiment analysis)은 ‘대상에 대한 대중의 긍부정 평가’이기 때문에 평가 대상과 맥락 요소가 어느 정도 한정되므로 담화에서 평가하기의 언어적 실현을 탐구하는 것이 가능해진다. 본 연구는 상품평 담화의 감성분석이라는 목적 아래 한국어 담화에서 대상에 대한 평가가 어떤 언어적 특질을 보이며 실현되는가에 대해 탐구한 연구이다. 연구의 대상은 용언, 부사, 명사를 포함한 모든 품사와 구, 절 단위까지 포괄하였으며 연구의 목적은 상품평 담화에서 어휘와 문법의 어떤 의미가 긍부정 평가를 나타내거나 혹은 조정할 수 있는지를 밝히는 데 두었다. 2장에서는 이론적 논의를 주제별로 감성분석과 감정분석 연구 분야 중 데이터 구축과 관련한 국내외 성과, 국어학과 언어 정보학의 어휘 분류 연구에 대하여 살펴보았다. 3장에서는 본 연구의 감성분석용 말뭉치의 구성과 분석 방법에 대해 기술하였다. 분석용 말뭉치는 국립국어원의 '속성 기반 감성분석 말뭉치 2021'를 기본 토대로 구성하였으며 부족분에 대해서는 네이버 및 트립어드바이저 상품평을 추가 수집하였다. 말뭉치는 총 82,289 어절로서 미용제품, 가전제품, 숙박업소, 영화 이렇게 총 4개의 도메인으로 구성되었다. 분석 방법은 먼저 4개 도메인 대상으로 감성 분류 체계를 수립하고 감성분석을 수행한 후 개체와 속성을 구분하였다. 그다음에 ‘속성내포형’과 ‘속성명시형’의 개념을 정립 후 감성표현을 이 두 유형별로 제시하였다. 속성내포형은 속성이 언어적으로 드러나지 않고 그 의미가 어휘에 포함된 형태이며 속성명시형은 속성이 언어적으로 명시된 형태를 가리킨다. 이 두 분류를 기초로 하여 4장에서 각 도메인별 속성 분석 결과와 속성별 감성분석 결과를 제시하였다. 4장을 통해 총 4개 도메인의 감성 분류 체계의 구성 예시와 각 도메인별 속성 표현, 감성표현 목록, 속성과 감성의 구분 문제에 대해 고찰해 볼 수 있었다. 미용·가전·숙박업소 도메인에서는 가격, 디자인, 사용감, 서비스 등 공통 평가 항목이 존재했으나, 영화는 유명도와 이용대상 외에는 겹치는 평가 항목이 거의 없었고 감정분석까지 함께 수행해야 평가가 완전해짐을 확인했다. 또한 속성 분석을 통해 동일한 속성 표현이 함께 결합하는 평가 표현에 따라 지시 속성이 달라질 수 있다는 점 등을 추가로 확인할 수 있었다. 5장에서는 어휘 표현을 의미 특성별로 유형화하여 이 의미 특성이 상품평 담화에서의 긍부정 평가 판별에 어떤 상관성이 있는지를 분석하였다. 상품평 담화에서 항상 긍부정 평가를 나타낼 수 있는 어휘 평가 표현은 호감도 표현, 감각 표현이 있었다. 맥락 조건 혹은 형태·통사적 조건이 주어질 때만 제한적으로 긍부정 평가를 보이는 어휘 평가 표현으로는 외형 묘사 표현, 성질 묘사 표현, 감정 표현, 지각인지 표현, 행위성 표현이 있었다. 마지막으로 비교 표현, 정도성 표현, 소재 표현은 감성표현을 구성하는 기능어로서의 역할을 담당하고 있었다. 6장에서는 문법 표현의 의미가 상품평 담화에서 긍부정 평가 실현에 어떤 영향을 미치는지를 분석하였다. 상품평 담화에서 항상 긍부정 평가를 보이는 문법 평가 표현은 긍부정 의미를 내포하거나 담화 내 쓰임이 긍부정 판단의 묘사인 경우였는데 ‘-(으)ㄹ 만하다’, ‘-어/아 보세요’, 가능의 ‘-(으)ㄹ 수 있다’, ‘-(으)면 되다’, ‘-어/아도 되다’. ‘-기는 하다’가 해당하였다. 둘째로 담화에서 특정한 어휘 평가 표현과 결합할 때 제한적으로 긍부정 평가를 보이는 문법 표현들로는 [조건], [소망], [의지]와 같이 미경험 전제의 문법 표현, 조사 ‘만큼, 만, 이다’, [명령]. [후회], 사동, [의무], ‘-어/아 버리다’, ‘-어/아 보이다’, [비유] 문법 표현이 있었다. 그리고 [목적], [시간] 문법 표현은 단순 만족도 표현과 같이 전형적으로 긍부정 평가를 드러내는 어휘 평가 표현과 결합 시 오히려 어휘의 긍부정 평가를 상실시키는 효과를 냈다. 본고의 연구 결과는 향후 다방면에서 감성분석용 지식 구축 시 필수적으로 고려해야 하는 감성 분류 체계, 속성 표현, 감성표현, 속성과 감성의 구분 문제에 대해 실제적으로 도움을 줄 수 있을 것이다. 또한 한국어 상품평 담화에서 긍부정 평가 판단에 영향을 미치는 어휘와 문법의 의미 특성을 규명하고 그 목록과 고려 요인을 제공하였다는 점에서 언어학 연구 성과를 실제 감성분석에 접목한 충분한 가치를 지닌다.
텍스트마이닝 기법을 활용한 레스토랑 고객의 감성분석에 관한 연구 : 외래관광객의 온라인 리뷰 빅데이터 중심으로
Along with an advent of the 4th industrial revolution, paradigms in various domestic industries have changed, where in the field of ‘hospitality’ is experiencing diverse changes based on digital technologies comprising big data and artificial intelligence etc. In particular, consumers are writing and sharing their experiences of services they had received real-time in accordance with increasing access to internet via the popularization of smartphones by which the data thereof are increasing explosively. Such a vast amount of review data can play the important role over entire domain of consumers’ decision making, as well as the roles of delivery of information and recommenders wherein the positive reviews on services they had experienced increase the loyalty of consumers. In addition, the review data reflect diverse kinds of consumers’ emotions frankly expressed thus they can be exploited as resources to measure and improve consumers’ service experiences. A survey employing questionnaire is widely used to measure customers’ experiences through which diverse kinds of questions can be used for quantification of consumers’ experiences however it has disadvantages of measurement errors dependent on limitations of sampling and errors in memory of past experiences of respondents, terminologies in questions, and length of questions etc. On the contrary, the analysis using text mining techniques has an advantage of identifying psychological aspects of respondents realized by the broad collection of all written opinions in categories freely set by consumers which was difficult to identify through conducting quantitative studies despite disadvantages of selecting appropriate respondents and difficulty in measurement according to intended designs of the analysis. The present study intends to read the emotional factors of satisfaction and dissatisfaction of foreign tourists by using the left on-line reviews of restaurants, and to verify the presence of positive effects of sentiment polarity values of selection attributes upon overall level of customers’ satisfaction and differences in types of restaurants. To conduct the study, the data of on-line reviews on hotels and restaurants of customers were collected from the Review Community (TripAdvisor). The crawler was prepared by using Python to collect the data from corresponding community by which the profile information of customers, reviews written in English, and comprehensive scores thereof were collected. And then, the text mining techniques were exploited for the analysis of each text. The natural language processing technique was used to quantify the text data, and the frequency analysis on keywords per each type of restaurants was carried out. And then, the topic modeling technique was used to distinguish each sentence in reviews in terms of selection attributes of the foods, price, service, and atmosphere, and then the distinguished sentences were put into the sentiment analysis to determine the factors of satisfaction or dissatisfaction of tourists. Finally, the test of hypotheses was carried out by using techniques of statistical analysis on the effectual relationship between differences in sentiment polarity at each type of restaurants and degree of satisfaction of customers. The results obtained from the present study are summarized as in the following. First, the factors of satisfaction and dissatisfaction per each type of restaurants were compared to each other through conducting sentiment analysis on the on-line reviews left by tourists. Common terms used to express dissatisfaction of tourists comprised time, staff, wait, late, forget, missing, and wrong etc. The terms imply the degree of hospitality of employees in restaurants and management of waiting time of customers in restaurants could influence greatly on the degree of dissatisfaction of customers. In respect of each type of restaurants, the terms comprising expensive, quality, course, table, reservation, credit, card, glass, plate, table, lobby, cat, and taxi etc. appeared in fine dining restaurants, whereas the terms such as English, speak, language, queue, and busy etc. appeared in casual dining restaurants. The terms such as line, quick, fast, seating, parking, wrap, and dirty were frequently used in describing the quick and self-service restaurants. This suggests the factors of dissatisfaction vary according to each type of restaurants. Second, the demographic characteristics of customers were compared by using the results of sentiment analysis. Contrary to customers visiting restaurants who left many reviews on respective services and foods, the reviews on prices appeared relatively few, wherein the difference in scores of emotion on selection attributes between sex and age of customers appeared insignificant. Besides, the ratio of dissatisfaction on price and service appeared higher than that of foods and atmosphere of restaurants from the analysis of customers divided into groups of the satisfaction (positive) and dissatisfaction (negative) of selection attributes. This suggests the service of restaurants requires continuous awareness and management in that the failure in providing pertinent service could lead to the dissatisfaction of customers. Third, the differences in values of sentiment polarity of selection attributes per each type of restaurants were compared. The value of sentiment polarity of selection attributes of foods, service, and atmosphere of fine dining restaurant appeared highest, whereas the highest value of sentiment polarity of selection attribute of price appeared in the quick and self-service restaurant. In respect of the appraisal of tourists on the high-class fine dining restaurants, they appeared satisfactory on foods, service, and atmosphere therein while they exhibited dissatisfaction on prices of restaurants. Fourth, the ANOVA tests per each type of restaurants were carried out to determine the presence of significant differences in values of sentiment polarity of selection attributes. The results revealed the presence of significant differences of all selection attributes per each type of restaurants. Fifth, the effect of sentiment polarity of selection attributes per each type of restaurants upon degree of satisfaction of customers was examined. Values of sentiment polarity of foods, price, service, and atmosphere of restaurants upon customers’ satisfaction were verified by conducting the Tobit regression analysis. The values of sentiment polarity of food, price, and service appeared with significant effect on customer’s satisfaction, whereas the effect of the value of sentiment polarity of atmosphere of restaurants appeared insignificant. This was concluded the preference to foods, price, and service than atmosphere of restaurants of practical tourists affected relatively more on the degrees of satisfaction. In the present study, the factors of satisfaction and dissatisfaction of tourists reflected in on-line reviews were analyzed and compared to each other by using the sentiment analysis of selection attributes per each type of restaurants based on the topic modeling technique, by which the study demonstrated the quantification of customers’ experiences was enabled through the information of text. Together with ordinary approaches that measured customers’ experiences by exploiting conventional surveys employing respective questionnaires, the study approach to text review also demonstrated that it could bring significant research results. Keywords selection attribute of restaurant, on-line review, big data, text mining, natural language processing, topic modeling, sentiment analysis 4차산업혁명이 도래하면서 국내 다양한 산업들의 패러다임이 변화하고 있고, Hospitality 분야에서도 빅데이터, 인공지능 등 디지털 기술을 기반으로 다양한 변화가 생기고 있다. 특히 스마트폰의 대중화 및 인터넷 접근성의 증가와 함께 소비자들은 서비스 경험을 실시간으로 온라인상 작성 및 공유하고 있고, 그 데이터의 양이 폭발적으로 증가하고 있다. 이러한 방대한 리뷰 데이터는 소비자 의사결정 전 영역에 걸쳐 중요한 영향력을 행사하고 있고, 정보 전달자와 추천인의 역할을 하는 한편, 긍정적으로 작성된 리뷰는 소비자의 충성도를 높이는 역할을 하고 있다. 또한 리뷰 데이터는 고객의 솔직하고 다양한 감성을 담고 있어 고객 서비스 경험을 측정하고 개선하는 소중한 자원으로 활용될 수 있다. 고객 경험을 측정하기 위해 설문조사 방식을 많이 활용하고 있는데, 설문방식의 경우 연구목적에 따라 다양한 설문을 통해 계량화 할 수 있다는 장점이 있으나, 샘플링의 한계와 응답자가 설문에 응답하면서 과거 기억의 오류 및 설문 용어ㆍ길이에 따라 측정 오류가 발생할 수 있다는 단점이 있다. 반면에 텍스트마이닝을 통한 분석은 사전에 연구자가 원하는 목적대로 설계하여 응답자 선정 후 측정하기 힘들다는 단점이 있으나, 고객이 자유롭게 작성한 모든 범주의 의견을 폭넓게 수집하여, 양적 연구에서 포착하기 힘들었던 심리적인 부분까지 밝힐 수 있다는 장점이 있다. 본 연구의 목적은 외래관광객이 남긴 온라인 레스토랑 리뷰를 활용하여 만족과 불만족 감정요인을 읽어내고, 선택속성별 감성극성값이 레스토랑 유형별로 차이가 있는지, 전체 고객만족에 긍정적인 영향을 미치는지 검증하고자 한다. 연구 주제를 해결하기 위해 호텔ㆍ레스토랑 리뷰 커뮤니티(TripAdvisor)에서 온라인 리뷰데이터를 수집하였다. 해당 커뮤니티에서 데이터를 수집하기 위해 Python을 활용하여 크롤러를 제작 후 고객 프로필 정보, 영문텍스트 리뷰, 종합평점 정보를 수집하였다. 그리고 텍스트를 분석하기 위하여 텍스트 마이닝 기법을 활용하였다. 텍스트 데이터를 계량화하기 위하여 자연어 처리(NLP) 기법을 사용하였고, 레스토랑 유형별 핵심 키워드 빈도분석을 실시하였다. 다음으로 리뷰별 문장을 음식ㆍ가격ㆍ서비스ㆍ분위기 선택속성별로 구분하기 위하여 토픽모델링 기법을 활용하였고, 이렇게 분류된 문장을 감성분석을 실시하여 외래관광객의 만족ㆍ불만족 요인을 심층 분석하였다. 마지막으로 레스토랑 유형별 감성극성값의 차이 및 고객만족에 미치는 영향관계에 관해 통계 분석을 통한 가설 검증을 실시하였다. 연구결과를 정리하면 다음과 같다. 첫째, 외래관광객이 남긴 온라인 리뷰에 대하여 감성분석을 통하여 레스토랑 유형별로 만족요인과 불만족요인을 비교하였다. 레스토랑 공통 불만족 용어로 사용된 것은 time, staff, wait, late, forget, missing, wrong 등이다. 즉, 레스토랑 대기시간 관리와 종업원의 친절도가 고객 불만족에 큰 영향을 줄 수 있음을 나타낸다. 레스토랑 유형별로 보면 파인다이닝 불만족 용어로는 expensive, quality, course, table, reservation, credit, card, glass, plate, table, lobby, cat, taxi 등이 나타났고, 캐주얼다이닝의 경우는 english, speak, language, queue, busy 등이 나타났으며, 퀵ㆍ셀프서비스 레스토랑의 경우 line, quick, fast, seating, parking, wrap, dirty 등이 자주 사용되었다. 즉, 레스토랑 유형별로 주로 사용되는 불만족 요인 키워드가 차이가 있음을 알 수 있다. 둘째, 감성분석 결과를 활용하여 인구통계적 특성을 비교하였다. 레스토랑 방문 고객은 음식과 서비스에 대하여 많은 리뷰를 작성하는 반면 가격에 대해서는 그렇지 않은 것으로 나타났고, 선택속성별 감성점수는 성별과 연령에 따른 차이는 크지 않은 것으로 나타났다. 또한 선택속성별 만족(긍정) 그룹과 불만족(부정) 그룹으로 구분하여 불만족 그룹의 특성을 파악해 본 결과 가격과 서비스가 음식과 분위기에 비해 불만족 비율이 높은 것을 알 수 있다. 즉, 서비스 실패는 쉽게 불만족으로 연결될 수 있다는 점에서 지속 관리가 필요한 요소임을 알 수 있다. 셋째, 레스토랑 유형별로 선택속성별 감성극성값의 차이를 비교하였다. 음식, 서비스, 분위기 선택속성 감성극성값은 파인다이닝 레스토랑이 제일 높았고, 가격선택속성 감성극성값은 퀵ㆍ셀프서비스 레스토랑이 제일 높았다. 외래관광객이 고급 파인다이닝 레스토랑에 대한 평가시 음식, 서비스, 분위기는 대체적으로 만족하지만 비싼 가격으로 인해 가격에 대해서는 만족도가 낮은 것을 알 수 있다. 넷째, 레스토랑 유형별로 선택속성 감성극성값이 유의한 차이가 있는지 검증하기 위해 레스토랑 유형별 분산분석(ANOVA)을 실시하였다. 그 결과 레스토랑 유형별로 모든 선택속성이 유의한 차이를 보였다. 다섯째, 레스토랑 선택속성별 감성극성값이 고객만족에 영향을 미치는지 살펴보았다. 음식ㆍ가격ㆍ서비스ㆍ분위기 감성극성값이 고객만족에 영향을 미치는지 검증하기 위해 토빗-회귀분석을 실시하였다. 음식ㆍ가격ㆍ서비스 감성극성값은 고객만족도에 유의하게 영향을 미치는 것으로 나타났고, 분위기 감성극성값은 유의하지 않게 나타났다. 이는 실리를 중시여기는 외래관광객의 특성상 상대적으로 분위기 보다는 음식, 가격, 서비스 속성이 만족에 더 큰 영향을 미치기 때문으로 볼 수 있다. 본 연구는 토픽모델링에 기반한 선택속성별 감성분석을 활용하여 온라인 리뷰에 담긴 속성별 고객의 만족 및 불만족 요인을 심층적으로 분석하였고, 레스토랑 유형에 따라 비교한 연구로서, 텍스트 정보를 통해 고객의 경험을 계량화하여 파악할 수 있다는 것을 보여주었다. 기존 설문 조사를 통해 고객 경험을 측정하던 일반론적인 접근방식과 더불어 텍스트 리뷰를 통한 연구방식 또한 유의미한 결과를 가져올 수 있음을 보여주었다. 핵심어 : 레스토랑 선택속성, 온라인 리뷰, 빅데이터, 텍스트마이닝, 자연어 처리, 토픽모델링, 감성분석
Synthetic Data Generation Techniques using GPT-3 Model for Enhancing Imbalanced Sentiment Analysis
Suhaeni Cici 이화여자대학교 대학원 2024 국내박사
In the sentiment analysis field, the prevalence of imbalanced data has presented a significant challenge to achieving accurate classification performance, especially in the case of user-generated content. Usually, most sentiment analysis datasets have tended to over-represent popular sentiments while leaving minority sentiments underrepresented, resulting in classifiers that perform well on majority classes but poorly on minority ones. Previous approaches to addressing this imbalance, such as the use of Generative Adversarial Networks (GANs) for synthetic data generation, have shown potential but also limitations, indicating a need for more advanced and effective methods. This dissertation introduces an innovative approach to overcoming this challenge by leveraging the GPT-3 model, a state-of-the-art language model known for its advanced ability to generate human-like text across a vast array of applications. Recognized for its deep learning capabilities and extensive training on diverse data, GPT-3 offers a promising solution for generating synthetic data that can balance sentiment classes effectively. This study focused on two primary objectives: first, to generate good-quality synthetic data to balance the training dataset in imbalanced sentiment analysis scenarios using two distinct strategies—fine-tuning the GPT-3 model (GPT-FT) and employing a sentence-by-sentence generation technique (GPT-SS); and second, to improve sentiment classification performance with a balanced dataset using sophisticated deep learning models, including RNN, CNN, LSTM, BiLSTM, and GRU, with GloVe embeddings for feature extraction. The methodology employed in this study involved generating synthetic data using the GPT-3 model to address the specific needs of underrepresented classes within the Coursera review dataset, a widely recognized source of user-generated content in educational sentiment analysis. This synthetic data was then evaluated for its quality based on novelty, diversity, and the relevance with the Coursera review context. The results indicated that while both GPT-FT and GPT-SS methods were effective in enhancing classification performance, GPT-FT demonstrated a better ability to generate more novel and diverse synthetic data. However, GPT-FT also faced challenges such as the occasional generation of anomalous data, highlighting areas for further refinement. Conversely, GPT-SS was able to generate synthetic data without anomalies but encountered problems with high similarity among the generated data. Additionally, the evaluation of sentiment classification showed that GPT-FT achieved higher performance compared to GPT-SS. Conclusively, this dissertation demonstrated that synthetic data augmentation using GPT-3 can significantly improve sentiment analysis performance by addressing the imbalance in class distribution. The proposed methods, GPT-FT and GPT-SS, not only contributed to the field of sentiment analysis by providing more balanced datasets but also by offering insights into the effective generation and evaluation of synthetic data. The study suggests further research into more sophisticated anomaly detection and data similarity management techniques to enhance the quality of synthetic data. Overall, the research establishes a foundational approach for using advanced language models like GPT-3 to tackle prevalent issues in sentiment analysis and opens new avenues for exploring the application of synthetic data in other areas of text analytics. 감정 분석 분야에서, 데이터의 불균형은 특히 사용자 생성 콘텐츠의 경우, 정확한 분류 성능을 달성하는 데 중대한 도전이 되었다. 일반적으로 대부분의 감정 분석 데이터셋은 인기 있는 감정을 과대표현하는 경향이 있으며 소수 감정은 상대적으로 부족하게 표현되어, 다수 클래스에 대해서는 잘 작동하지만 소수 클래스에 대해서는 성능이 떨어지는 분류기를 만들게 된다. 이러한 불균형을 해결하기 위한 이전의 접근법들, 예를 들어 합성 데이터 생성을 위한 생성적 적대 신경망(GANs)의 사용은 잠재력을 보였지만 한계도 있음을 나타내며, 더 발전된 효과적인 방법이 필요함을 시사했다. 이논문은 GPT-3 모델을 활용하여 이러한 도전을 극복하는 혁신적인 접근법을 소개한다. GPT-3는 다양한 애플리케이션에서 인간과 같은 텍스트를 생성할 수 있는 탁월한 능력으로 알려진 최신 언어 모델이다. 깊은 학습 능력과 다양한 데이터에 대한 광범위한 훈련을 통해 인정받은 GPT-3는 감정 클래스를 효과적으로 균형 잡히게 할 수 있는 합성 데이터를 생성하는 유망한 해결책을 제공하였다. 이 연구는 두 가지 주요 목표에 중점을 두었다. 첫째, 불균형 감정 분석 시나리오에서 훈련 데이터셋을 균형잡히게 하기 위해 GPT-3 모델을 미세 조정하는 방법(GPT-FT)과 문장별 생성 기술(GPT-SS)을 사용하여 양질의 합성 데이터를 생성하는 것이다. 둘째, GloVe 임베딩을 특징 추출에 사용하면서 RNN, CNN, LSTM, BiLSTM, GRU 등 고급 딥러닝 모델을 사용하여 균형 잡힌 데이터셋으로 감정 분류 성능을 향상시키는 것이다. 이 연구에서 사용한 방법론은 교육 감정 분석에서 사용자 생성 콘텐츠의 널리 알려진 출처인 Coursera 리뷰 데이터셋 내 소수 클래스의 특정 요구를 해결하기 위해 GPT-3 모델을 사용하여 합성 데이터를 생성하는 것을 포함했다. 이 합성 데이터는 그 후 신규성, 다양성 및 일관성을 기준으로 그 품질이 평가되었다. 결과는 GPT-FT와 GPT-SS 방법 모두 분류 성능을 향상시키는 데 효과적이었지만, GPT-FT가 더욱 새롭고 다양한 합성 데이터를 생성하는 능력이 더 뛰어났음을 보여주었다. 그러나 때때로 이상 데이터 생성과 같은 문제도 발생하여 추가 개선이 필요한 영역을 강조했다. 반면에 GPT-SS는 이상 데이터 없이 합성 데이터를 생성할 수 있었지만 생성된 데이터 간에 높은 유사성을 보이는 문제가 있었다. 또한 감정 분류 평가는 GPT-FT가 GPT-SS에 비해 더 높은 성능을 달성했음을 보여주었다. 결론적으로, 이 논문은 GPT-3를 사용한 합성 데이터 증강이 클래스 분포의 불균형을 해결함으로써 감정 분석 성능을 크게 향상시킬 수 있음을 입증했다. 제안된 방법들, GPT-FT 및 GPT-SS는 감정 분석 분야에 더 균형 잡힌 데이터셋을 제공함으로써 기여할 뿐만 아니라, 합성 데이터의 효과적인 생성 및 평가에 대한 통찰을 제공한다. 연구는 합성 데이터의 품질을 향상시키기 위해 더 정교한 이상 감지 및 데이터 유사성 관리 기술에 대한 추가 연구를 제안한다. 전반적으로, 이 연구는 감정 분석에서 널리 있는 문제를 해결하기 위해 GPT-3와 같은 고급 언어 모델을 사용하는 기초적인 접근 방식을 확립하고, 다른 텍스트 분석 분야에서 합성 데이터의 활용을 탐구할 새로운 길을 열었다.
이지연 과학기술연합대학원대학교 2023 국내박사
Recently, along with eco-friendly issues, the importance of the global carbon material industry is increasing. Demand for lightweight and high-strength carbon materials is increasing in terms of weight reduction of materials, which are the key to eco-friendliness. Among the six core carbon material technologies, although graphene has created a strong wave of enthusiasm after its discovery in 2004 and the subsequent Nobel prize awarded to the two detectors in 2010, there has been a lot of skepticism about the level of its commercialization in the industry and the public. Under such the gap between academia and industry, there is a lack of research on how graphene research and development (R&D) trend has changed, and how people around the world are accepting its potential. Therefore, this study intends to analyze major issues and public sentiments of graphene from 2010 to 2021 using text mining methods. To this end, this research presents a method of analyzing technology trends from a new perspective by converging national R&D project information, newspaper articles, and social media data. Focusing on the fact that hype phenomenon appears in relation to society in promising new technologies, the hype cycle theory is used to comprehensively compare and analyze each result. Through this, this study aims to present a complementary technology trend analysis method that can be used when establishing the direction of technology management strategy and national R&D policy by identifying the correlation between R&D and media from a hype cycle perspective. To examine how far graphene technology R&D has come, a total of 5,693 projects from 2010 to 2021 were collected from the National Science and Technology Information Service (NTIS), which is National Science and Technology Information Service of South Korea. Text mining techniques such as keyword frequency analysis, association rule mining, and topic modelling were applied to the collected national R&D project information. As a result, graphene R&D investment showed an increasing trend from 2010 to 2016, focusing on basic research, and then started to decrease after 2017. However, the proportion of applied research and development research began to increase in 2019, which can be seen that the investment to enter the practical stage of graphene technology development is expanding. Secondly, the public sentiment was investigated in order to identify sentiment shift and major discussion issues related to graphene in the media. Media data included 1,103 electronic newspaper articles and 11,287 comments from global social media Reddit. Through media, the overall trend and expectations of the industry and the public on graphene have been analyzed by using sentiment analysis. This study found that over the half of the comments were positive toward graphene, and specific sentiments fluctuated commensurate with major events, for example, Nobel Prize, commercialization delay, and wide applications to the industry. At the final step of the analysis, by examining whether the results derived from each data correspond to the trajectory of the hype cycle, it was demonstrated that R&D, media, and public opinion on social media exchange the influence under the hype phenomenon. This study provides a convincing case study for understanding shifts in the sentiments of the interested public along with policy, R&D, and media, which will interest graphene technology R&D experts and policy-makers. This paper contributes to the understanding of how the South Korean government has been investing in graphene technology R&D, as well as global graphene sentiments, and therefore providing reference for strategic research planning and policy making. The text mining method that converges heterogeneous data based on the hype cycle presented in this research can be applied to analyze other emerging technology trends and market expectations. It is expected that these data will be used as reference materials for medium and long-term technology management plans and related policy establishment. 최근 친환경 이슈와 함께 글로벌 탄소소재 산업에 대한 중요성이 더해지고 있다. 친환경의 핵심인 소재의 경량화 측면에서 가볍고 강도가 높은 탄소소재의 수요가 증가하고 있다. 탄소소재 6대 핵심 기술 중 하나인 그래핀은 2004년 영국 맨체스터 대학교의 두 교수에 의해 발견되어 2010년에 노벨상을 수상하게 되면서 뜨거운 관심을 불러일으켰지만, 그동안 산업계나 대중들 사이에서는 상용화 수준에 대한 회의적인 논의가 많았다. 학계와 산업계 간 격차가 존재하는 상황에서 그래핀 연구개발(R&D) 동향이 어떻게 변화해 왔으며 전 세계 사람들이 그 가능성에 대해 어떻게 받아들이고 있는지에 대한 연구는 부족하다. 따라서 본 연구에서는 텍스트마이닝을 활용하여 2010년부터 2021년까지의 그래핀 주요 이슈 및 대중 감성을 분석하고자 한다. 이를 위해 국가R&D 과제정보, 신문기사, 소셜미디어 댓글 데이터를 혼용하여 새로운 시각에서의 기술 동향 분석 방법을 제시하고, 유망 신기술은 사회와의 관계에서 하이프 현상이 나타난다는 점에 착안하여 하이프 사이클 이론을 활용하여 각 결과들을 종합적으로 비교 분석한다. 이를 통해 본 연구는 R&D와 미디어 간의 상관관계를 하이프 사이클 관점에서 파악하여 국가R&D 정책 및 전략적 기술경영의 방향성을 수립할 때 활용할 수 있는 보완적인 기술 동향 분석 방법을 제시하는 것을 목표로 한다. 먼저 우리나라 그래핀 R&D의 투자 동향을 파악하기 위해 국가과학기술지식정보서비스(NTIS)에서 다운받은 총 5,693개의 과제 정보를 연구 대상으로 삼았다. 수집된 국가R&D 과제정보에 대해 정량적 현황 분석을 수행하고, 키워드 빈도 분석, 연관규칙 분석, 토픽모델링 등의 텍스트마이닝 기법을 적용한 정성적 분석을 더하였다. 그 결과, 그래핀 R&D 투자는 기초연구 위주로 2010년부터 2016년까지 증가하다가 그 이후로 감소하는 추세를 보인 반면 응용연구와 개발연구 비중이 2019년을 기점으로 증가하기 시작하였다. 이러한 결과로 보아 그래핀은 기술 개발의 실용화 단계에 접어들기 위한 투자가 확대되고 있음을 알 수 있었다. 두 번째로는 미디어에서 그래핀과 관련된 주요 키워드와 대중 감성을 파악해보았다. 미디어는 전자 매체와 소셜미디어로 구분하여 각각 국내 과학기술 분야 일간 전문지인 전자신문에서 1,103건, 글로벌 소셜미디어인 레딧(Reddit)에서 11,287건의 데이터를 수집하였다. 전자 매체를 통해서는 주요 키워드 및 주제 분류를 통해 전체적인 동향과 기사 내용에서 나타난 감정의 전반적 극성을 파악하고, 소셜미디어를 통해서는 그래핀에 관심이 있는 대중의 감성 변화와 주요 논의 주제를 식별하였다. 마지막으로 각 데이터를 통해 도출된 결과들이 하이프 사이클의 궤적과 부합하는지 여부를 살펴봄으로써 정책, R&D, 미디어 및 여론이 서로 상관관계를 가지고 하이프의 영향을 주고받는 특징을 밝혀냈다. 본 연구는 그래핀과 같은 신기술과 관련하여 R&D 전문가와 정책 입안자가 정책 의제 설정 등에 활용할 수 있는 설득력 있는 사례를 제공한다. 국가R&D 과제정보와 신문기사의 텍스트마이닝을 통해 그래핀 R&D 영역을 탐색하고, 기술 동향 분석을 위한 소셜미디어 데이터의 감성분석이라는 새로운 접근을 통해 글로벌 감성분석 결과의 국내 활용 가능성을 검토해보았다. 본 논문에서 제시한 하이프 사이클 기반의 이종 데이터 혼용 텍스트마이닝 방법은 신흥 기술 트렌드 및 시장의 기대치 분석에 활용하여 기술의 중장기 경영 계획 및 관련 정책 수립에 참고 자료로 활용할 수 있을 것으로 기대된다.
The Construction of a Korean Pre-Trained Model and an Enhanced Application on Sentiment Analysis
최근 트랜스포머 양방향 인코더 표현 (Bidirectional Encoder Representations from Transformers, BERT) 모델에 대한 관심이 높아지면서 자연어처리 분야에서 이에 기반한 연구 역시 활발히 이루어지고 있다. 이러한 문장 단위의 임베딩을 위한 모델들은 보통 학습 과정에서 문장 내 어휘, 통사, 의미 정보를 포착하여 모델링한다고 알려져 있다. 따라서 ELMo, GPT, BERT 등은 그 자체가 다양한 자연어처리 문제를 해결할 수 있는 보편적인 모델로서 기능한다. 본 연구는 한국어 자료로 학습한 단일 언어 BERT 모델을 제안한다. 가장 먼저 공개된 한국어를 다룰 수 있는 BERT 모델은 Google Research의 multilingual BERT (M-BERT)였다. 이는 한국어와 영어를 포함하여 104개 언어로 구성된 학습 데이터와 어휘 목록을 가지고 학습한 모델이며, 모델 하나로 포함된 모든 언어의 텍스트를 처리할 수 있다. 그러나 이는 그 다중언어성이 갖는 장점에도 불구하고, 각 언어의 특성을 충분히 반영하지 못하여 단일 언어 모델보다 각 언어의 텍스트 처리 성능이 낮다는 단점을 보인다. 본 연구는 그러한 단점들을 완화하면서 텍스트에 포함되어 있는 언어 정보를 보다 잘 포착할 수 있도록 구성된 데이터와 어휘 목록을 이용하여 모델을 구축하고자 하였다. 따라서 본 연구에서는 한국어 Wikipedia 텍스트와 뉴스 기사로 구성된 데이터를 이용하여 KR-BERT 모델을 구현하고, 이를 GitHub을 통해 공개하여 한국어 정보처리를 위해 사용될 수 있도록 하였다. 또한 해당 학습 데이터에 댓글 데이터와 법조문과 판결문을 덧붙여 확장한 텍스트에 기반해서 다시 KR-BERT-MEDIUM 모델을 학습하였다. 이 모델은 해당 학습 데이터로부터 WordPiece 알고리즘을 이용해 구성한 한글 중심의 토큰 목록을 사전으로 이용하였다. 이들 모델은 개체명 인식, 질의응답, 문장 유사도 판단, 감정 분석 등의 다양한 한국어 자연어처리 문제에 적용되어 우수한 성능을 보고했다. 또한 본 연구에서는 BERT 모델에 감정 자질을 추가하여 그것이 감정 분석에 특화된 모델로서 확장된 기능을 하도록 하였다. 감정 자질을 포함하여 별도의 임베딩 모델을 학습시켰는데, 이때 감정 자질은 문장 내의 각 토큰에 한국어 감정 분석 코퍼스 (KOSAC)에 대응하는 감정 극성(polarity)과 강도(intensity) 값을 부여한 것이다. 각 토큰에 부여된 자질은 그 자체로 극성 임베딩과 강도 임베딩을 구성하고, BERT가 기본으로 하는 토큰 임베딩에 더해진다. 이렇게 만들어진 임베딩을 학습한 것이 감정 자질 모델(sentiment-combined model)이 된다. KR-BERT와 같은 학습 데이터와 모델 구성을 유지하면서 감정 자질을 결합한 모델인 KR-BERT-KOSAC를 구현하고, 이를 GitHub을 통해 배포하였다. 또한 그로부터 학습 과정 내 언어 모델링과 감정 분석 과제에서의 성능을 얻은 뒤 KR-BERT와 비교하여 감정 자질 추가의 효과를 살펴보았다. 또한 감정 자질 중 극성과 강도 값을 각각 적용한 모델을 별도 구성하여 각 자질이 모델 성능 향상에 얼마나 기여하는지도 확인하였다. 이를 통해 두 가지 감정 자질을 모두 추가한 경우에, 그렇지 않은 다른 모델들에 비하여 언어 모델링이나 감정 분석 문제에서 성능이 어느 정도 향상되는 것을 관찰할 수 있었다. 이때 감정 분석 문제로는 영화평의 긍부정 여부 분류와 댓글의 악플 여부 분류를 포함하였다. 그런데 위와 같은 임베딩 모델을 사전학습하는 것은 많은 시간과 하드웨어 등의 자원을 요구한다. 따라서 본 연구에서는 비교적 적은 시간과 자원을 사용하는 간단한 모델 결합 방법을 제시한다. 적은 수의 인코더 레이어, 어텐션 헤드, 적은 임베딩 차원 수로 구성한 감정 자질 모델을 적은 스텝 수까지만 학습하고, 이를 기존에 큰 규모로 사전학습되어 있는 임베딩 모델과 결합한다. 기존의 사전학습모델에는 충분한 언어 모델링을 통해 다양한 언어 처리 문제를 처리할 수 있는 보편적인 기능이 기대되므로, 이러한 결합은 서로 다른 장점을 갖는 두 모델이 상호작용하여 더 우수한 자연어처리 능력을 갖도록 할 것이다. 본 연구에서는 감정 분석 문제들에 대한 실험을 통해 두 가지 모델의 결합이 학습 시간에 있어 효율적이면서도, 감정 자질을 더하지 않은 모델보다 더 정확한 예측을 할 수 있다는 것을 확인하였다. Recently, as interest in the Bidirectional Encoder Representations from Transformers (BERT) model has increased, many studies have also been actively conducted in Natural Language Processing based on the model. Such sentence-level contextualized embedding models are generally known to capture and model lexical, syntactic, and semantic information in sentences during training. Therefore, such models, including ELMo, GPT, and BERT, function as a universal model that can impressively perform a wide range of NLP tasks. This study proposes a monolingual BERT model trained based on Korean texts. The first released BERT model that can handle the Korean language was Google Research’s multilingual BERT (M-BERT), which was constructed with training data and a vocabulary composed of 104 languages, including Korean and English, and can handle the text of any language contained in the single model. However, despite the advantages of multilingualism, this model does not fully reflect each language’s characteristics, so that its text processing performance in each language is lower than that of a monolingual model. While mitigating those shortcomings, we built monolingual models using the training data and a vocabulary organized to better capture Korean texts’ linguistic knowledge. Therefore, in this study, a model named KR-BERT was built using training data composed of Korean Wikipedia text and news articles, and was released through GitHub so that it could be used for processing Korean texts. Additionally, we trained a KR-BERT-MEDIUM model based on expanded data by adding comments and legal texts to the training data of KR-BERT. Each model used a list of tokens composed mainly of Hangul characters as its vocabulary, organized using WordPiece algorithms based on the corresponding training data. These models reported competent performances in various Korean NLP tasks such as Named Entity Recognition, Question Answering, Semantic Textual Similarity, and Sentiment Analysis. In addition, we added sentiment features to the BERT model to specialize it to better function in sentiment analysis. We constructed a sentiment-combined model including sentiment features, where the features consist of polarity and intensity values assigned to each token in the training data corresponding to that of Korean Sentiment Analysis Corpus (KOSAC). The sentiment features assigned to each token compose polarity and intensity embeddings and are infused to the basic BERT input embeddings. The sentiment-combined model is constructed by training the BERT model with these embeddings. We trained a model named KR-BERT-KOSAC that contains sentiment features while maintaining the same training data, vocabulary, and model configurations as KR-BERT and distributed it through GitHub. Then we analyzed the effects of using sentiment features in comparison to KR-BERT by observing their performance in language modeling during the training process and sentiment analysis tasks. Additionally, we determined how much each of the polarity and intensity features contributes to improving the model performance by separately organizing a model that utilizes each of the features, respectively. We obtained some increase in language modeling and sentiment analysis performances by using both the sentiment features, compared to other models with different feature composition. Here, we included the problems of binary positivity classification of movie reviews and hate speech detection on offensive comments as the sentiment analysis tasks. On the other hand, training these embedding models requires a lot of training time and hardware resources. Therefore, this study proposes a simple model fusing method that requires relatively little time. We trained a smaller-scaled sentiment-combined model consisting of a smaller number of encoder layers and attention heads and smaller hidden sizes for a few steps, combining it with an existing pre-trained BERT model. Since those pre-trained models are expected to function universally to handle various NLP problems based on good language modeling, this combination will allow two models with different advantages to interact and have better text processing capabilities. In this study, experiments on sentiment analysis problems have confirmed that combining the two models is efficient in training time and usage of hardware resources, while it can produce more accurate predictions than single models that do not include sentiment features.
Khushboo Shafi 고려대학교 대학원 2024 국내석사
This thesis investigates the optimization of Airbnb operations in Singapore through a detailed analysis of guest preferences using K-means clustering and aspect-based sentiment analysis. The primary objectives of the study are to understand guest preferences, identify key aspects influencing guest satisfaction, and provide actionable insights for Airbnb hosts. By systematically analyzing guest reviews of Airbnb listings in Singapore, the study employs K-means clustering to categorize listings into distinct clusters based on similarities in their features, thus revealing the diverse needs and preferences of different guest segments. The research identifies and evaluates specific aspects of Airbnb listings, such as accuracy, amenities, check-in process, comfort, communication, family-friendliness, host interaction, location, maintenance and condition, neighborhood, and value, which significantly impact guest satisfaction. By pinpointing the most influential aspects, hosts can allocate resources more effectively and address critical areas needing enhancement. The findings are translated into practical recommendations that hosts can implement to optimize their listings and improve guest satisfaction, guiding them in making informed decisions that enhance their operational strategies, leading to better guest reviews and increased bookings. Key outcomes of this analysis include ensuring accuracy in listing descriptions to align with guest expectations, enhancing the check-in process with clear instructions and flexible options, improving communication through timely responses and proactive engagement, offering good value by competitively pricing listings and highlighting unique features, focusing on host interaction by being approachable and providing personalized touches, and investing in comfort and amenities to maintain and upgrade quality furnishings and essential amenities regularly. Furthermore, hosts will gain a strategic framework to optimize their listings based on identified guest preferences and sentiment trends, which includes personalizing guest experiences, leveraging positive aspects highlighted by guests, and creating unique selling propositions. Insights from this study contribute to a broader understanding of guest preferences and behaviors within the Singaporean Airbnb market, providing data-driven evidence to support policy decisions regarding short-term rentals, offering benchmarks for other hosts to evaluate and improve their listings, and encouraging higher standards of service and quality across the Airbnb community. 본 논문은 K-평균 군집화와 Aspect-Based Sentiment Analysis 를 사용하여 싱가포르의 에어비앤비 운영을 최적화하는 방법을 탐구한다. 연구의 주요 목적은 고객 선호도를 이해하고, 고객 만족도에 영향을 미치는 주요 측면을 식별하며, 에어비앤비 호스트에게 실질적인 인사이트를 제공하는 것이다. 싱가포르 에어비앤비 숙소를 대상으로 K-평균 군집화를 통해 유사한 특성을 가진 숙소를 군집화하고 고객 리뷰를 체계적으로 분석하여 다양한 고객 세그먼트에 있어 니즈와 선호도를 밝혀내었다. 본 연구에서는 정확성, 편의시설, 체크인 절차, 편안함, 소통, 가족 친화성, 호스트 상호작용, 위치, 유지보수 상태, 이웃 환경, 그리고 가치를 포함한 에어비앤비 숙소의 특징을 식별하고 평가하여 고객 만족도에 큰 영향을 미치는 요소를 파악한다. 가장 영향력있는 측면을 정확히 파악함으로써 호스트는 자원을 더 효과적으로 할당하고 개선이 필요한 핵심 영역의 문제를 해결할 수 있다. 이러한 발견은 호스트가 숙소를 최적화하고 고객 만족도를 높이기 위해 구현할 수 있는 실질적인 실행 아이템으로 정리되며, 운영 전략을 강화하여 개선된 고객 리뷰 및 예약 증가를 이끌어낼 수 있다. 호스트에 대한 이 연구의 주요 메세지는 고객 기대 수준에 맞게 숙소 설명을 정확히 하고 신속한 응답과 적극적인 소통을 통해 커뮤니케이션을 향상시킬 것, 명확한 설명과 유연한 선택사항으로 체크인 절차를 개선할 것, 경쟁력 있는 가격 책정과 차별화된 기능을 통해 가치를 제공할 것, 접근이 용이하고 개인화된 접촉을 제공하여 호스트와의 상호작용을 효율화할 것, 고품질의 가구와 필수 편의시설을 정기적으로 유지하고 업그레이드하여 편안함을 갖춘 숙소가 되도록 지속적으로 편의시설에 투자할 것 등이다. 호스트는 식별된 고객 선호도와 감성 트렌드에 따라 숙소를 최적화하기 위한 전략적 프레임워크를 얻을 수 있으며, 이는 개인화된 고객 경험과 고객의 강조 측점을 활용하고, 차별화된 판매 포인트를 창출하는 것을 포함한다. 본 연구의 인사이트를 요약하면, 싱가포르 에어비앤비 시장 내 고객 선호도와 행동에 대한 더 넓은 이해를 돕고, 단기 임대에 관한 정책 결정을 지원하는 데이터 기반의 근거를 제공하며, 다른 호스트가 자신의 숙소를 평가하고 개선할 수 있는 벤치마크를 제공하고, 에어비앤비 커뮤니티 전반에 걸쳐 서비스와 품질의 요구 수준을 파악하게 해준다.