언어 모델은 발전해왔지만, 국방 용어 이해에 특화된 한국어 모델은 부족하다. 기존 연구는 주로 BERT 기반 한국어 사전학습에 집중했으나, 본 논문은 다국어 사전학습된 BGE-M3를 국방 분야에 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=A110167788
2026
Korean
embedding model ; contrastive learning ; hard negative ; BGE-M3 ; 임베딩 모델 ; 대조학습 ; Hard Negative ; BGE-M3
KCI우수등재
학술저널
117-123(7쪽)
0
상세조회0
다운로드언어 모델은 발전해왔지만, 국방 용어 이해에 특화된 한국어 모델은 부족하다. 기존 연구는 주로 BERT 기반 한국어 사전학습에 집중했으나, 본 논문은 다국어 사전학습된 BGE-M3를 국방 분야에 ...
언어 모델은 발전해왔지만, 국방 용어 이해에 특화된 한국어 모델은 부족하다. 기존 연구는 주로 BERT 기반 한국어 사전학습에 집중했으나, 본 논문은 다국어 사전학습된 BGE-M3를 국방 분야에 맞게 미세조정(Fine-tuning)하고, 대조학습(Contrastive Learning) 시 네거티브 샘플 선택이 성능에 미치는 영향을 분석했다. 네거티브 샘플링 방식으로 (1) 무작위(Easy), (2) 구분이 어려운(Hard), (3) 유사도가 높은(Harder) 방식을 비교한 결과, Harder 샘플링이 Accuracy, NMI, ARI에서 가장 우수했다. 또한, KorSTS 데이터셋으로 미세조정 전후의 Pearson 및 Spearman 상관계수를 분석한 결과, 기존 모델의 언어적 지식을 효과적으로 유지하면서도 국방 도메인에 특화됨을 확인했다.
다국어 초록 (Multilingual Abstract)
Korean language models specifically designed for the defense sector are still limited, even with the rapid advancements in text embeddings. In this study, we fine-tune the multilingual BGE-M3 model to better understand military terminology and investi...
Korean language models specifically designed for the defense sector are still limited, even with the rapid advancements in text embeddings. In this study, we fine-tune the multilingual BGE-M3 model to better understand military terminology and investigate how negative sampling in contrastive learning impacts downstream performance. We evaluate three strategies: Easy (random negatives), Hard (lexicographic adjacency), and Harder (similarity-mined negatives). Our analysis, based on clustering metrics such as Accuracy, NMI, and ARI using a defense news dataset, reveals that the similarity-based Harder strategy consistently outperforms the others. Further evaluations on the KorSTS dataset demonstrate that the Harder approach maintains strong Spearman and Pearson correlations, indicating successful domain adaptation without compromising overall semantic competence. Interestingly, the three Harder variants—negatives mined with BGE-M3, ko-sroberta, and multilingual-e5—produce nearly identical similarity distributions and comparable improvements, while the Easy strategy plateaus and the Hard strategy shows only moderate performance. These findings suggest that mining sufficiently similar negatives, as opposed to using random or adjacent ones, is crucial for effective, domain-specific fine-tuning of multilingual embedding models.
VR 헤드셋 사용자의 블렌드쉐입 기반 표정 트래킹을 통한 파라메트릭 모델 및 사실적인 아바타 재구성
Q 함수 기반 Lyapunov 안정성 제약을 통한 강화 학습 알고리즘의 안정성과 성능 향상
전이학습기반 심층 강화학습 알고리즘을 활용한 동적 환경에서의 공간 적응적 자율이동 탐색 무인기 설계