RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    K-MIR: A Korean Multimodal Dataset for Intent Recognition and Lightweight Model Optimization = K-MIR: 한국어 멀티모달 의도 인식 데이터셋 구축 및 경량 모델 최적화 연구

    한글로보기

    https://www.riss.kr/link?id=T17282558

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent advances in natural language processing (NLP) and computer vision (CV) have accelerated research on multimodal intent recognition, aiming to integrate text, audio, and video modalities to better understand human intentions. However, existing multimodal intent recognition datasets are predominantly English-based, limited in scale, and often fail to capture the complexity of real-world conversational flows. Particularly, there is a significant lack of multimodal datasets applicable to Korean, hindering the practical expansion of Korean intent recognition technologies. In this study, we constructed K-MIR, a Korean multimodal intent recognition dataset, by benchmarking the English-based MIntRec dataset and utilizing real-world broadcast content. K-MIR comprises 2,270 utterances across 20 intent classes and includes text, audio, and video modalities. For data construction and feature extraction, tools such as Whisper, PySceneDetect, and YOLOv8 were employed, and labeling was conducted using a majority voting scheme to ensure reliability. To evaluate the quality of K-MIR, we conducted experiments using the MAG-BERT model, which not only confirmed that multimodal fusion yields significant performance improvements over text-only inputs, but also demonstrated that the same model architecture performs stably and achieves comparable benchmark-level results on K-MIR, thereby validating its practical utility and reliability. In parallel, to address the structural complexity and computational inefficiency of existing Transformer-based models such as MAG-BERT and MULT, this study proposes three lightweight multimodal classification models: FusionBERT, GatedFusionNet, and AlignContrastiveNet. These models were integrated using a 3-Way Soft Voting ensemble strategy as a practical approach to enhancing performance. To enable fair comparison with previous studies and assess the generalizability of the proposed architectures, training and evaluation were conducted using the publicly available MIntRec benchmark dataset, and the results consistently outperformed baseline models. In particular, the proposed hybrid ensemble model achieved an accuracy of 74% and a macro-F1 score of 0.69, surpassing the performance of MAG-BERT. In addition, the contrastive learning-based hybrid structure was designed to align semantically similar cross-modal representations within a shared embedding space, while the Soft Voting ensemble effectively leveraged the complementary strengths of individual models to further enhance overall performance. This study made three key contributions: providing a practical data resource (K-MIR) for Korean-language multimodal intent recognition, developing lightweight model architectures with enhanced performance, and optimizing modality fusion strategies. These contributions establish a solid foundation for advancing the practicality and future scalability of Korean multimodal AI systems.
    번역하기

    Recent advances in natural language processing (NLP) and computer vision (CV) have accelerated research on multimodal intent recognition, aiming to integrate text, audio, and video modalities to better understand human intentions. However, existing mu...

    Recent advances in natural language processing (NLP) and computer vision (CV) have accelerated research on multimodal intent recognition, aiming to integrate text, audio, and video modalities to better understand human intentions. However, existing multimodal intent recognition datasets are predominantly English-based, limited in scale, and often fail to capture the complexity of real-world conversational flows. Particularly, there is a significant lack of multimodal datasets applicable to Korean, hindering the practical expansion of Korean intent recognition technologies. In this study, we constructed K-MIR, a Korean multimodal intent recognition dataset, by benchmarking the English-based MIntRec dataset and utilizing real-world broadcast content. K-MIR comprises 2,270 utterances across 20 intent classes and includes text, audio, and video modalities. For data construction and feature extraction, tools such as Whisper, PySceneDetect, and YOLOv8 were employed, and labeling was conducted using a majority voting scheme to ensure reliability. To evaluate the quality of K-MIR, we conducted experiments using the MAG-BERT model, which not only confirmed that multimodal fusion yields significant performance improvements over text-only inputs, but also demonstrated that the same model architecture performs stably and achieves comparable benchmark-level results on K-MIR, thereby validating its practical utility and reliability. In parallel, to address the structural complexity and computational inefficiency of existing Transformer-based models such as MAG-BERT and MULT, this study proposes three lightweight multimodal classification models: FusionBERT, GatedFusionNet, and AlignContrastiveNet. These models were integrated using a 3-Way Soft Voting ensemble strategy as a practical approach to enhancing performance. To enable fair comparison with previous studies and assess the generalizability of the proposed architectures, training and evaluation were conducted using the publicly available MIntRec benchmark dataset, and the results consistently outperformed baseline models. In particular, the proposed hybrid ensemble model achieved an accuracy of 74% and a macro-F1 score of 0.69, surpassing the performance of MAG-BERT. In addition, the contrastive learning-based hybrid structure was designed to align semantically similar cross-modal representations within a shared embedding space, while the Soft Voting ensemble effectively leveraged the complementary strengths of individual models to further enhance overall performance. This study made three key contributions: providing a practical data resource (K-MIR) for Korean-language multimodal intent recognition, developing lightweight model architectures with enhanced performance, and optimizing modality fusion strategies. These contributions establish a solid foundation for advancing the practicality and future scalability of Korean multimodal AI systems.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 자연어처리와 컴퓨터비전 기술의 발전으로 텍스트, 오디오, 비디오 등 다양한 모달리티를 통합하여 인간의 의도를 보다 정확하게 이해하는 멀티모달 의도 인식 연구가 활발히 이루어지고 있다. 그러나 멀티모달 의도 인식 데이터셋은 대부분 영어 기반이며, 스케일이 제한적이고 실제 대화 흐름을 충분히 반영하지 못한다. 특히 한국어 환경에 적용 가능한 멀티모달 데이터셋은 부족하여, 한국어 기반 의도 인식 기술의 실질적 확장에 장애가 되고 있다. 본 연구에서는 영어권 MIntRec 데이터셋을 벤치마크로 삼아 실제 방송 콘텐츠를 활용하여 한국어 환경에 적합한 멀티모달 의도 인식 데이터셋인 K-MIR를 구축하였다. K-MIR는 텍스트, 오디오, 비디오 모달을 모두 포함하며, 총 2,270개의 발화와 20개 의도 클래스를 포괄한다. 데이터 구성 및 특징 추출 과정에서는 Whisper, PySceneDetect, YOLOv8 등을 활용하였으며, 라벨링은 다수결 원칙에 따라 신뢰성을 확보하였다. 또한 K-MIR의 품질을 평가하기 위해 MAG-BERT 모델을 적용한 실험을 수행한 결과, 멀티모달 융합이 텍스트 단독 입력 대비 의미 있는 성능 향상을 가져옴을 확인하였을 뿐 아니라, 동일 모델 구조가 K-MIR 에서도 안정적으로 작동하며 기존 벤치마크 수준의 성능을 재현함으로써, 본 데이터셋의 실용성과 신뢰성 또한 입증할 수 있었다. 한편, 기존 Transformer 기반 모델(MAG-BERT, MULT 등)의 구조적 복잡성과 연산량 문제를 극복하고자, 본 연구에서는 FusionBERT, GatedFusionNet, AlignContrastiveNet 세 가지 경량 멀티모달 분류 모델을 설계하고, 이들을 3-Way Soft Voting 기반으로 앙상블하여 실용적인 성능 향상 전략을 구성하였다. 제안된 모델은 기존 연구들과의 공정한 성능 비교 및 아키텍처의 일반화 가능성 평가를 위해, 공용 벤치마크인 MIntRec 데이터셋을 기반으로 학습 및 검증을 수행하였으며, 그 결과 baseline 모델 대비 일관된 성능 우위를 보였다. 특히, 제안된 하이브리드 앙상블 모델은 Accuracy 74%, Macro-F1 0.69 를 기록하며, MAG-BERT 대비 우수한 성능을 달성하였다. 이와 함께, 대조 학습 기반의 하이브리드 구조는 의미적으로 유사한 모달리티 간 표현이 공통 임베딩 공간에서 가깝게 정렬되도록 설계되었으며, Soft Voting 앙상블은 개별 모델의 상호보완적 특성을 극대화하여 전체 성능 향상에 기여하였다. 본 연구는 한국어 기반 멀티모달 인식 기술을 위한 실질적인 데이터 자원(K-MIR) 제공, 모델 구조의 경량화 및 성능 개선을 달성하였으며, 모달리티 융합 전략을 최적화함으로써 한국어 멀티모달 AI 시스템의 실용성과 연구 확장 가능성을 뒷받침할 수 있는 기반을 마련하였다.
    번역하기

    최근 자연어처리와 컴퓨터비전 기술의 발전으로 텍스트, 오디오, 비디오 등 다양한 모달리티를 통합하여 인간의 의도를 보다 정확하게 이해하는 멀티모달 의도 인식 연구가 활발히 이루어...

    최근 자연어처리와 컴퓨터비전 기술의 발전으로 텍스트, 오디오, 비디오 등 다양한 모달리티를 통합하여 인간의 의도를 보다 정확하게 이해하는 멀티모달 의도 인식 연구가 활발히 이루어지고 있다. 그러나 멀티모달 의도 인식 데이터셋은 대부분 영어 기반이며, 스케일이 제한적이고 실제 대화 흐름을 충분히 반영하지 못한다. 특히 한국어 환경에 적용 가능한 멀티모달 데이터셋은 부족하여, 한국어 기반 의도 인식 기술의 실질적 확장에 장애가 되고 있다. 본 연구에서는 영어권 MIntRec 데이터셋을 벤치마크로 삼아 실제 방송 콘텐츠를 활용하여 한국어 환경에 적합한 멀티모달 의도 인식 데이터셋인 K-MIR를 구축하였다. K-MIR는 텍스트, 오디오, 비디오 모달을 모두 포함하며, 총 2,270개의 발화와 20개 의도 클래스를 포괄한다. 데이터 구성 및 특징 추출 과정에서는 Whisper, PySceneDetect, YOLOv8 등을 활용하였으며, 라벨링은 다수결 원칙에 따라 신뢰성을 확보하였다. 또한 K-MIR의 품질을 평가하기 위해 MAG-BERT 모델을 적용한 실험을 수행한 결과, 멀티모달 융합이 텍스트 단독 입력 대비 의미 있는 성능 향상을 가져옴을 확인하였을 뿐 아니라, 동일 모델 구조가 K-MIR 에서도 안정적으로 작동하며 기존 벤치마크 수준의 성능을 재현함으로써, 본 데이터셋의 실용성과 신뢰성 또한 입증할 수 있었다. 한편, 기존 Transformer 기반 모델(MAG-BERT, MULT 등)의 구조적 복잡성과 연산량 문제를 극복하고자, 본 연구에서는 FusionBERT, GatedFusionNet, AlignContrastiveNet 세 가지 경량 멀티모달 분류 모델을 설계하고, 이들을 3-Way Soft Voting 기반으로 앙상블하여 실용적인 성능 향상 전략을 구성하였다. 제안된 모델은 기존 연구들과의 공정한 성능 비교 및 아키텍처의 일반화 가능성 평가를 위해, 공용 벤치마크인 MIntRec 데이터셋을 기반으로 학습 및 검증을 수행하였으며, 그 결과 baseline 모델 대비 일관된 성능 우위를 보였다. 특히, 제안된 하이브리드 앙상블 모델은 Accuracy 74%, Macro-F1 0.69 를 기록하며, MAG-BERT 대비 우수한 성능을 달성하였다. 이와 함께, 대조 학습 기반의 하이브리드 구조는 의미적으로 유사한 모달리티 간 표현이 공통 임베딩 공간에서 가깝게 정렬되도록 설계되었으며, Soft Voting 앙상블은 개별 모델의 상호보완적 특성을 극대화하여 전체 성능 향상에 기여하였다. 본 연구는 한국어 기반 멀티모달 인식 기술을 위한 실질적인 데이터 자원(K-MIR) 제공, 모델 구조의 경량화 및 성능 개선을 달성하였으며, 모달리티 융합 전략을 최적화함으로써 한국어 멀티모달 AI 시스템의 실용성과 연구 확장 가능성을 뒷받침할 수 있는 기반을 마련하였다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Preface v
    • Contents vi
    • List of Tables ix
    • List of Figures x
    • Abstract i
    • Preface v
    • Contents vi
    • List of Tables ix
    • List of Figures x
    • 1 Introduction 1
    • 1.1 Background and Motivation 1
    • 1.2 Research Objectives 2
    • 1.3 Thesis Structure 3
    • 2 Related Work 5
    • 2.1 Development of Intent Recognition 5
    • 2.2 Multimodal Learning and Intent Recognition 6
    • 2.3 Representative Multimodal Datasets 6
    • 2.4 Advances in Multimodal Model Architectures 7
    • 2.5 Limitations and Contributions of This Study 9
    • 3 K-MIR Dataset Construction 11
    • 3.1 Data Sources and Construction Procedures 11
    • 3.2 Annotation and Speaker Information Processing 14
    • 3.3 Multimodal Feature Extraction 16
    • 3.4 Dataset Statistics and Composition 18
    • 3.5 Dataset Vaidation and Release Plan 19
    • 4 Proposed Method 25
    • 4.1 FusionBERT 25
    • 4.2 GatedFusionNet 28
    • 4.3 AlignContrastiveNet 30
    • 4.4 3-Way Soft Voting Ensemble 32
    • 4.5 Comparison with MAG-BERT 34
    • 5 Experiment Settings 36
    • 5.1 Objectives and Strategy 36
    • 5.2 Dataset: MIntRec 37
    • 5.3 Input Processing and Feature Construction 39
    • 5.4 Model Configurations 40
    • 5.5 Hyperparameter Settings 41
    • 5.6 Evaluation Metrics 42
    • 5.7 Experimmental Environment 43
    • 6 Experimental Results and Analysis 44
    • 6.1 Performance Comparison 44
    • 6.2 Ablation Study Results 45
    • 6.3 Summary of Experimental Findings 47
    • 7 Conclusion and Future Work 48
    • 7.1 Research Summary and Key Contributions 48
    • 7.2 Insights from Experimental Results 49
    • 7.3 Future Research Directions 51
    • 7.4 Final Remarks 52
    • References 53
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼