본 연구는 대규모 텍스트를 대상으로 하는 토픽 모델링에서 반복 실행에 따라 결과가 달라지는 구조적 불안정성이 해석 가능성과 재현성을 저해할 수 있다는 문제의식에서 출발하였다. 지...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17372197
서울 : 국민대학교 일반대학원, 2025
학위논문(석사) -- 국민대학교 일반대학원 , 데이터사이언스전공 , 2026. 2
2025
한국어
결과안정성(Result stability) ; BERTopic ; LDA ; S_CLOP ; Topic modeling ; Result stability ; BERTopic ; LDA ; S-CLOP ; Topic modeling
서울
vii, 87 ; 26 cm
지도교수: 정여진
I804:11014-200000957235
0
상세조회0
다운로드본 연구는 대규모 텍스트를 대상으로 하는 토픽 모델링에서 반복 실행에 따라 결과가 달라지는 구조적 불안정성이 해석 가능성과 재현성을 저해할 수 있다는 문제의식에서 출발하였다. 지...
본 연구는 대규모 텍스트를 대상으로 하는 토픽 모델링에서 반복 실행에 따라 결과가 달라지는 구조적 불안정성이 해석 가능성과 재현성을 저해할 수 있다는 문제의식에서 출발하였다. 지금까지의 안정성 관련 논의는 주로 LDA 계열의 확률 기반 모델에 집중되어 있었으며, BERTopic과 같은 임베딩 기반 최신 모델에 대한 구조적 안정성과 실행 간 일관성의 체계적인 검토는 상대적으로 부족하였다.
이에 본 연구는 LDA Prototype의 반복 실행 및 대표 실행 선택 전략을 임베딩 기반 BERTopic 모델에 적용한 Prototype-BERTopic을 제안한다. 본 연구의 실험 진행은 기본 LDA, BERTopic, LDA-Prototype, 제안 모델인 Prototype-BERTopic을 대상으로, Coherence, Diversity, Significance, Pairwise Topic Similarity, S-CLOP, S̄ 등의 다양한 정량적 지표를 활용하여 성능을 비교하였다.
실험 결과, Prototype-BERTopic은 반복된 실행 사이의 구조적 일관성과 재현성 측면에서 기존 BERTopic 및 LDA 기반 모델을 평균적으로 상회하였으며 특히, 기존의 Coherence 지표가 임베딩 기반 모델에 불리하게 작용할 수 있다는 해석상의 한계도 함께 분석하여, 평가 지표를 선택하는 것의 중요성을 강조하였다.
본 연구는 임베딩 기반 토픽 모델의 구조적 안정성과 실행 신뢰성 문제를 실험적을 통해 검증하고, 이를 보완할 수 있는 실용적인 대안 모델을 제시함으로써 연구의 재현성과 실 적용 가능성의 동시 향상을 나타내고자 한다.
다국어 초록 (Multilingual Abstract)
This study addresses the issue of structural instability in topic modeling for large-scale text corpora, where repeated executions under the same conditions may yield inconsistent results, thereby hindering interpretability and reproducibility. While ...
This study addresses the issue of structural instability in topic modeling for large-scale text corpora, where repeated executions under the same conditions may yield inconsistent results, thereby hindering interpretability and reproducibility. While previous discussions on model stability have largely focused on probabilistic models such as LDA, systematic evaluations of embedding-based models—notably BERTopic—remain limited.
To bridge this gap, we propose Prototype-BERTopic, which adapts the representative execution selection strategy from LDA Prototype to the BERTopic framework. We conduct experiments using standard LDA, BERTopic, Prototype LDA, and Prototype-BERTopic evaluating their performance based on multiple quantitative metrics, including Coherence, Diversity, Significance, Pairwise Topic Similarity, S-CLOP, and Rn Score.
Our experimental results demonstrate that Prototype-BERTopic consistently outperforms baseline models in terms of structural consistency and reproducibility across executions. Moreover, BERTopic-ILR effectively reduces topic redundancy while enhancing representational diversity. We also highlight potential limitations of traditional metrics—such as Coherence—when applied to embedding-based approaches, emphasizing the importance of metric selection in model evaluation. By empirically validating the structural reliability of embedding-based topic models and proposing practical alternatives, this study contributes to improving both the reproducibility and applicability of topic modeling in real-world text analysis tasks.evaluating their performance based on multiple quantitative metrics, including Coherence, Diversity, Significance, Pairwise Topic Similarity, S-CLOP, and Rn Score. Our experimental results demonstrate that Prototype-BERTopic consistently outperforms baseline models in terms of structural consistency and reproducibility across executions.
Moreover, BERTopic-ILR effectively reduces topic redundancy while enhancing representational diversity. We also highlight potential limitations of traditional metrics—such as Coherence—when applied to embedding-based approaches, emphasizing the importance of metric selection in model evaluation. By empirically validating the structural reliability of embedding-based topic models and proposing practical alternatives, this study contributes to improving both the reproducibility and applicability of topic modeling in real-world text analysis tasks.
목차 (Table of Contents)