RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI우수등재

    국방 언어 임베딩 모델을 위한 BGE-M3의 미세조정: 대조학습에서 네거티브 샘플 선택이 모델 성능에 미치는 영향 = Fine-Tuning BGE-M3 for Defense Language Embedding Model: The Impact of Negative Sample Selection in Contrastive Learning

    한글로보기

    https://www.riss.kr/link?id=A110167788

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    언어 모델은 발전해왔지만, 국방 용어 이해에 특화된 한국어 모델은 부족하다. 기존 연구는 주로 BERT 기반 한국어 사전학습에 집중했으나, 본 논문은 다국어 사전학습된 BGE-M3를 국방 분야에 맞게 미세조정(Fine-tuning)하고, 대조학습(Contrastive Learning) 시 네거티브 샘플 선택이 성능에 미치는 영향을 분석했다. 네거티브 샘플링 방식으로 (1) 무작위(Easy), (2) 구분이 어려운(Hard), (3) 유사도가 높은(Harder) 방식을 비교한 결과, Harder 샘플링이 Accuracy, NMI, ARI에서 가장 우수했다. 또한, KorSTS 데이터셋으로 미세조정 전후의 Pearson 및 Spearman 상관계수를 분석한 결과, 기존 모델의 언어적 지식을 효과적으로 유지하면서도 국방 도메인에 특화됨을 확인했다.
    번역하기

    언어 모델은 발전해왔지만, 국방 용어 이해에 특화된 한국어 모델은 부족하다. 기존 연구는 주로 BERT 기반 한국어 사전학습에 집중했으나, 본 논문은 다국어 사전학습된 BGE-M3를 국방 분야에 ...

    언어 모델은 발전해왔지만, 국방 용어 이해에 특화된 한국어 모델은 부족하다. 기존 연구는 주로 BERT 기반 한국어 사전학습에 집중했으나, 본 논문은 다국어 사전학습된 BGE-M3를 국방 분야에 맞게 미세조정(Fine-tuning)하고, 대조학습(Contrastive Learning) 시 네거티브 샘플 선택이 성능에 미치는 영향을 분석했다. 네거티브 샘플링 방식으로 (1) 무작위(Easy), (2) 구분이 어려운(Hard), (3) 유사도가 높은(Harder) 방식을 비교한 결과, Harder 샘플링이 Accuracy, NMI, ARI에서 가장 우수했다. 또한, KorSTS 데이터셋으로 미세조정 전후의 Pearson 및 Spearman 상관계수를 분석한 결과, 기존 모델의 언어적 지식을 효과적으로 유지하면서도 국방 도메인에 특화됨을 확인했다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Korean language models specifically designed for the defense sector are still limited, even with the rapid advancements in text embeddings. In this study, we fine-tune the multilingual BGE-M3 model to better understand military terminology and investigate how negative sampling in contrastive learning impacts downstream performance. We evaluate three strategies: Easy (random negatives), Hard (lexicographic adjacency), and Harder (similarity-mined negatives). Our analysis, based on clustering metrics such as Accuracy, NMI, and ARI using a defense news dataset, reveals that the similarity-based Harder strategy consistently outperforms the others. Further evaluations on the KorSTS dataset demonstrate that the Harder approach maintains strong Spearman and Pearson correlations, indicating successful domain adaptation without compromising overall semantic competence. Interestingly, the three Harder variants—negatives mined with BGE-M3, ko-sroberta, and multilingual-e5—produce nearly identical similarity distributions and comparable improvements, while the Easy strategy plateaus and the Hard strategy shows only moderate performance. These findings suggest that mining sufficiently similar negatives, as opposed to using random or adjacent ones, is crucial for effective, domain-specific fine-tuning of multilingual embedding models.
    번역하기

    Korean language models specifically designed for the defense sector are still limited, even with the rapid advancements in text embeddings. In this study, we fine-tune the multilingual BGE-M3 model to better understand military terminology and investi...

    Korean language models specifically designed for the defense sector are still limited, even with the rapid advancements in text embeddings. In this study, we fine-tune the multilingual BGE-M3 model to better understand military terminology and investigate how negative sampling in contrastive learning impacts downstream performance. We evaluate three strategies: Easy (random negatives), Hard (lexicographic adjacency), and Harder (similarity-mined negatives). Our analysis, based on clustering metrics such as Accuracy, NMI, and ARI using a defense news dataset, reveals that the similarity-based Harder strategy consistently outperforms the others. Further evaluations on the KorSTS dataset demonstrate that the Harder approach maintains strong Spearman and Pearson correlations, indicating successful domain adaptation without compromising overall semantic competence. Interestingly, the three Harder variants—negatives mined with BGE-M3, ko-sroberta, and multilingual-e5—produce nearly identical similarity distributions and comparable improvements, while the Easy strategy plateaus and the Hard strategy shows only moderate performance. These findings suggest that mining sufficiently similar negatives, as opposed to using random or adjacent ones, is crucial for effective, domain-specific fine-tuning of multilingual embedding models.

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼