RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    소셜 미디어 데이터의 감정 분석 : 사전 학습된 BERT를 이용한 이중 미세 조정 접근 = Emotion Analysis in Social Media Data: A Dual Fine-Tuning Approach Using Pre-trained BERT

    한글로보기

    https://www.riss.kr/link?id=T16984745

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    With the advent of the 21st century, advancements in the IT industry, alongside the proliferation of smartphones and the internet, have brought significant changes to people's daily lives and consumption of culture. Particularly, platforms like YouTube have disrupted the traditional paradigm of broadcast content production, opening new realms for individual creators and small-scale content production. These shifts have led to a wide consumption of diverse contents alongside the innovation of streaming services and the emergence of OTT (Over-The-Top) platforms, enabling both corporations and individual creators to engage in active marketing and content creation.

    However, this modern social trend has not only benefitted OTT platforms. Notably, terms like 'poverty in abundance' and 'Netflix Syndrome' have emerged around Netflix, highlighting a phenomenon where the expansion of choices in movies and dramas induces users' decision-making time and mental stress. Additionally, the surge in small-scale content production on social network services like YouTube has increased the tendency of users to refer to relatively shorter review videos than the somewhat lengthy movies or dramas offered on OTT platforms.

    The popularity of small-scale content production and short review videos reflects users' emotions and preferences. Through emotion analysis, it's possible to understand these trends and establish more effective content production and marketing strategies.

    This study utilizes the BERT – BASE version of the BERT – Multilingual model to transcribe voices from videos containing YouTube creators' poetic interpretations into subtitles and experiments with emotion analysis through dual fine-tuning without pre-training, using comments from viewers. The emotion analysis execution mechanism is divided into two experiments: the first using comments and the second using subtitles, involving data exploration, preprocessing, model training, and performance evaluation based on accuracy.

    The study examines the correlation of labels using binary methods based on the emotion analysis of sampled subtitles and comments. For instance, subtitles of a 'cohabitation drama review' video predominantly showed 'jealousy' (Emotion Label, 31), while the comments reflected 'satisfaction' (Emotion Label, 54) and 'excitement' (Emotion Label, 55). This difference is attributed to factors like content and user response disparity, the diversity and subjectivity of emotions, and the dramatization and direction style. This natural variance between the emotions in drama subtitles and viewer comments illustrates how each individual's unique experiences, interpretations, and responses generate diverse emotional reactions.

    Particularly, performance evaluation results can be qualitatively compared with related studies. Despite the same learning environment, an increase in accuracy by 0.43% was proven, and BERT demonstrated somewhat higher performance through dual fine-tuning without special pre-training.
    번역하기

    With the advent of the 21st century, advancements in the IT industry, alongside the proliferation of smartphones and the internet, have brought significant changes to people's daily lives and consumption of culture. Particularly, platforms like YouTub...

    With the advent of the 21st century, advancements in the IT industry, alongside the proliferation of smartphones and the internet, have brought significant changes to people's daily lives and consumption of culture. Particularly, platforms like YouTube have disrupted the traditional paradigm of broadcast content production, opening new realms for individual creators and small-scale content production. These shifts have led to a wide consumption of diverse contents alongside the innovation of streaming services and the emergence of OTT (Over-The-Top) platforms, enabling both corporations and individual creators to engage in active marketing and content creation.

    However, this modern social trend has not only benefitted OTT platforms. Notably, terms like 'poverty in abundance' and 'Netflix Syndrome' have emerged around Netflix, highlighting a phenomenon where the expansion of choices in movies and dramas induces users' decision-making time and mental stress. Additionally, the surge in small-scale content production on social network services like YouTube has increased the tendency of users to refer to relatively shorter review videos than the somewhat lengthy movies or dramas offered on OTT platforms.

    The popularity of small-scale content production and short review videos reflects users' emotions and preferences. Through emotion analysis, it's possible to understand these trends and establish more effective content production and marketing strategies.

    This study utilizes the BERT – BASE version of the BERT – Multilingual model to transcribe voices from videos containing YouTube creators' poetic interpretations into subtitles and experiments with emotion analysis through dual fine-tuning without pre-training, using comments from viewers. The emotion analysis execution mechanism is divided into two experiments: the first using comments and the second using subtitles, involving data exploration, preprocessing, model training, and performance evaluation based on accuracy.

    The study examines the correlation of labels using binary methods based on the emotion analysis of sampled subtitles and comments. For instance, subtitles of a 'cohabitation drama review' video predominantly showed 'jealousy' (Emotion Label, 31), while the comments reflected 'satisfaction' (Emotion Label, 54) and 'excitement' (Emotion Label, 55). This difference is attributed to factors like content and user response disparity, the diversity and subjectivity of emotions, and the dramatization and direction style. This natural variance between the emotions in drama subtitles and viewer comments illustrates how each individual's unique experiences, interpretations, and responses generate diverse emotional reactions.

    Particularly, performance evaluation results can be qualitatively compared with related studies. Despite the same learning environment, an increase in accuracy by 0.43% was proven, and BERT demonstrated somewhat higher performance through dual fine-tuning without special pre-training.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    21세기 들어 IT 산업의 발전과 함께, 스마트폰과 인터넷 보급은 사람들의 일상과 문화 소비 방식에 커다란 변화를 가져왔다. 특히 유튜브와 같은 플랫폼의 등장은 전통적인 방송 콘텐츠 제작 패러다임을 깨고, 개인 크리에이터와 소규모 콘텐츠 제작에 참여할 수 있는 새로운 영역을 열었다. 이러한 변화는 폭넓은 스트리밍 서비스의 혁신과 OTT(Over – The –Top, 이하 생략) 플랫폼의 등장과 함께 다양한 콘텐츠가 소비 됐으며, 이 콘텐츠를 바탕으로 기업과 개인 크리에이터들은 활발한 마케팅 및 창작물을 제공했다.

    그러나 이러한 현대의 사회적 흐름은 OTT 플랫폼들에게 이점만 남기진 않았다. 특히, 넷플릭스를 중심으로 ‘풍요 속의 빈곤’과 ‘넷플릭스 증후군’이라는 용어들이 등장함을 통해 영화나 드라마의 선택의 폭을 넓힘으로써 사용자의 선택 시간과 정신적인 스트레스를 유발하는 현상이 나타났다. 더불어, 유튜브와 같은 소셜 네트워크 서비스를 대상으로 소규모 콘텐츠 제작의 열풍을 일으키며, 다소 긴 영화나 드라마를 제공하는 OTT 플랫폼의 영화나 드라마들보다 비교적 짧은 리뷰 영상을 참고하는 사용자의 경향이 높아지고 있다.

    이러한 소규모 콘텐츠 제작의 열풍과 짧은 리뷰 영상의 인기는 사용자의 감정과 선호도를 반영하는데, 감정 분석을 통해 이러한 경향을 파악하고 콘텐츠 제작 및 마케팅 전략을 보다 효과적으로 수립할 수 있다.

    본 연구에서는 BERT – BASE 버전의 BERT – Multilingual 모델을 활용하여, 유튜브 크리에이터들의 서정적인 해석이 내포된 영상의 음성을 자막으로 텍스트화시키고, 해당 영상을 시청한 사용자들의 감정을 댓글을 활용하여 사전 학습 없이 이중 미세 조정을 통한 감정 분석을 실험한다. 감정 분석 실행 메커니즘은 댓글을 활용한 1차 실험과 자막을 활용한 2차 실험으로 나뉘어, 데이터 탐색, 전처리, 모델 학습, 정확도를 활용한 성능 평가로 이루어진다.

    연구 결과는 표본으로 추출한 자막의 감정 분석과 댓글의 감정 분석 결과를 토대로 이진법을 활용해 라벨의 연관성을 살펴보았으며, 대표적인 예로 간 떨어지는 동거 리뷰 영상의 자막은 주로 질투하는(감정라벨, 31번)이 나왔으며, 댓글 감정 분석 결과는 만족하는(감정라벨, 54번)과 흥분한(감정라벨, 55번)이 주로 나타났다. 이는 자막과 같은 감정이 나타나지 않았는데, 주요 큰 요인으로썬 콘텐츠와 사용자 반응의 차이, 감정의 다양성과 주관성, 드라마 표현 방식과 연출 때문이다. 이러한 요인들을 종합해 볼 때, 드라마의 자막 데이터와 시청자 댓글 사이의 감정의 차이가 나타나는 것은 매우 자연스러운 현상이다. 이는 각 개인의 독특한 경험, 해석 및 반응이 어떻게 다양한 감정적 반응을 생성하는지 보여준다.

    특히, 성능 평가 결과는 본연구와 관련 연구의 비교를 통해 정성적 성능 평가 결과를 살펴볼 수 있다. 이는 같은 학습 환경임에도 불구하고, 정확도는 0.43% 높음을 입증할 수 있었고, 특별한 사전 학습 없이도, BERT는 이중 미세 조정을 통해 다소 높은 성능을 보여줄 수 있음을 확인한다.
    번역하기

    21세기 들어 IT 산업의 발전과 함께, 스마트폰과 인터넷 보급은 사람들의 일상과 문화 소비 방식에 커다란 변화를 가져왔다. 특히 유튜브와 같은 플랫폼의 등장은 전통적인 방송 콘텐츠 제작 ...

    21세기 들어 IT 산업의 발전과 함께, 스마트폰과 인터넷 보급은 사람들의 일상과 문화 소비 방식에 커다란 변화를 가져왔다. 특히 유튜브와 같은 플랫폼의 등장은 전통적인 방송 콘텐츠 제작 패러다임을 깨고, 개인 크리에이터와 소규모 콘텐츠 제작에 참여할 수 있는 새로운 영역을 열었다. 이러한 변화는 폭넓은 스트리밍 서비스의 혁신과 OTT(Over – The –Top, 이하 생략) 플랫폼의 등장과 함께 다양한 콘텐츠가 소비 됐으며, 이 콘텐츠를 바탕으로 기업과 개인 크리에이터들은 활발한 마케팅 및 창작물을 제공했다.

    그러나 이러한 현대의 사회적 흐름은 OTT 플랫폼들에게 이점만 남기진 않았다. 특히, 넷플릭스를 중심으로 ‘풍요 속의 빈곤’과 ‘넷플릭스 증후군’이라는 용어들이 등장함을 통해 영화나 드라마의 선택의 폭을 넓힘으로써 사용자의 선택 시간과 정신적인 스트레스를 유발하는 현상이 나타났다. 더불어, 유튜브와 같은 소셜 네트워크 서비스를 대상으로 소규모 콘텐츠 제작의 열풍을 일으키며, 다소 긴 영화나 드라마를 제공하는 OTT 플랫폼의 영화나 드라마들보다 비교적 짧은 리뷰 영상을 참고하는 사용자의 경향이 높아지고 있다.

    이러한 소규모 콘텐츠 제작의 열풍과 짧은 리뷰 영상의 인기는 사용자의 감정과 선호도를 반영하는데, 감정 분석을 통해 이러한 경향을 파악하고 콘텐츠 제작 및 마케팅 전략을 보다 효과적으로 수립할 수 있다.

    본 연구에서는 BERT – BASE 버전의 BERT – Multilingual 모델을 활용하여, 유튜브 크리에이터들의 서정적인 해석이 내포된 영상의 음성을 자막으로 텍스트화시키고, 해당 영상을 시청한 사용자들의 감정을 댓글을 활용하여 사전 학습 없이 이중 미세 조정을 통한 감정 분석을 실험한다. 감정 분석 실행 메커니즘은 댓글을 활용한 1차 실험과 자막을 활용한 2차 실험으로 나뉘어, 데이터 탐색, 전처리, 모델 학습, 정확도를 활용한 성능 평가로 이루어진다.

    연구 결과는 표본으로 추출한 자막의 감정 분석과 댓글의 감정 분석 결과를 토대로 이진법을 활용해 라벨의 연관성을 살펴보았으며, 대표적인 예로 간 떨어지는 동거 리뷰 영상의 자막은 주로 질투하는(감정라벨, 31번)이 나왔으며, 댓글 감정 분석 결과는 만족하는(감정라벨, 54번)과 흥분한(감정라벨, 55번)이 주로 나타났다. 이는 자막과 같은 감정이 나타나지 않았는데, 주요 큰 요인으로썬 콘텐츠와 사용자 반응의 차이, 감정의 다양성과 주관성, 드라마 표현 방식과 연출 때문이다. 이러한 요인들을 종합해 볼 때, 드라마의 자막 데이터와 시청자 댓글 사이의 감정의 차이가 나타나는 것은 매우 자연스러운 현상이다. 이는 각 개인의 독특한 경험, 해석 및 반응이 어떻게 다양한 감정적 반응을 생성하는지 보여준다.

    특히, 성능 평가 결과는 본연구와 관련 연구의 비교를 통해 정성적 성능 평가 결과를 살펴볼 수 있다. 이는 같은 학습 환경임에도 불구하고, 정확도는 0.43% 높음을 입증할 수 있었고, 특별한 사전 학습 없이도, BERT는 이중 미세 조정을 통해 다소 높은 성능을 보여줄 수 있음을 확인한다.

    더보기

    목차 (Table of Contents)

    • 제1장 서론 = 1
    • 제2장 BERT 감정 분석 관련 연구 = 7
    • 제1절 BERT의 추론 과정 = 7
    • 제2절 트랜스포머 아키텍처 내부 = 9
    • 1. 트랜스포머 아키텍처 내부 = 9
    • 제1장 서론 = 1
    • 제2장 BERT 감정 분석 관련 연구 = 7
    • 제1절 BERT의 추론 과정 = 7
    • 제2절 트랜스포머 아키텍처 내부 = 9
    • 1. 트랜스포머 아키텍처 내부 = 9
    • 2. 트랜스포머 모델에서 단어 간의 정보 추출 방법 = 10
    • 3. 트랜스포머 모델의 전체적인 작동 원리 = 11
    • 4. 셀프 어텐션 메커니즘 속 스케일 닷 프로덕트 어텐션 계산 = 11
    • 5. 멀티 헤드 어텐션 = 12
    • 6. 피드 포워드 신경망 = 12
    • 7. 잔차 연결 = 14
    • 8. 레이어 정규화 = 14
    • 제3절 BERT의 표준적인 감정 분석 학습 단계 = 15
    • 제3장 BERT 이중 미세 조정 감정 분석 프로세스 설계 및 검증 전략 = 16
    • 제1절 이중 미세 조정 감정 분석 프로세스 구조 및 흐름 = 16
    • 제2절 이중 미세 조정 감정 분석 개요 및 검증 전략 = 19
    • 제4장 BERT 이중 미세 조정 감정 분석 실험 설계 및 구현 = 22
    • 제1절 BERT 이중 미세 조정 감정 분석 실험 프로세스 = 22
    • 1. 실험 데이터 소스 확인 = 22
    • 2. 실험 데이터 수집 = 23
    • 3. 댓글 데이터 전처리 및 탐색 = 26
    • (1) 환경 설정 = 26
    • (2) 1차 데이터 전처리 : 데이터 로드, 컬럼 추가, 병합 = 27
    • (3) 2차 데이터 전처리 : 데이터 정제를 위한, 결측치 제거와 타입 정의 = 28
    • (4) 1차 데이터 탐색 : 데이터 검증을 위한, 검토 및 파악 = 29
    • (5) 3차 데이터 전처리 = 29
    • (6) 2차 데이터 탐색 : 채널별 게시물 수, 드라마 별 게시물 수, 댓글 빈도 시각화 = 30
    • 4. 감정 모델 학습 준비 단계 : 감정 라벨 정의 및 학습 준비 = 34
    • 5. 댓글 데이터 감정 분석 학습 및 결과 = 37
    • (1) 모델 학습 : 첫 번째, 미세 조정 = 37
    • (2) 첫 번째, 미세 조정 모델 학습 결과 = 38
    • 6. 자막 데이터 전처리 및 탐색 = 39
    • (1) 환경 설정 : 글씨체 설정 및 필수 라이브러리 구축 = 39
    • (2) 1차 데이터 전처리 : 데이터 로드, 변환, 컬럼 추가, 병합 = 40
    • (3) 1차 데이터 탐색 : 데이터 타입 및 형식 검증, 무결성 검사 = 41
    • (4) 2차 데이터 탐색 : 드라마 줄거리 요약 채널의 주요 단어와 자막 분포 시각화 = 42
    • (5) 2차 데이터 전처리 : 불용어 제거, 토큰화 = 52
    • (6) 감성 분석 모델 학습 준비 단계 : 훈련/ 검증 데이터 분할, 모델 설정 = 53
    • 7. 자막 데이터 감정 분석 학습 및 결과 = 55
    • (1) 모델 학습 방법 = 55
    • (2) 모델 학습 결과 = 55
    • 8. 종합적 감정 분석 결과 = 56
    • (1) 표본 자막과 댓글의 감정 라벨 결과 = 56
    • (2) 이진 히트맵 = 57
    • (3) 이진 히트맵 시각화 목적 = 57
    • (4) 이진 히트맵 시각화 결과 = 61
    • 9. 모델 학습 평가 결과 = 61
    • 제5장 결론 = 63
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼