RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    한국어 특허 검색을 위한 장문 임베딩 모델 = A Study on Domain-Specific Long-Context Embedding Models for Advancing Korean Patent Retrieval

    한글로보기

    https://www.riss.kr/link?id=T17556256

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 전 세계적인 기술 경쟁 심화로 특허 출원이 지속적으로 증가함에 따라, 방대한 특허 문헌 속에서 필요한 정보를 효과적으로 탐색하고 분석하는 기술의 중요성이 대두되고 있다. 그러나 특허 분야에서 공개된 기존 언어모델은 주로 특허 말뭉치로 사전학습된 모델이며, 문서 간 유사도 검색을 위해 임베딩 공간을 최적화한 공개 모델은 부족한 실정이다. 또한 대부분 512 토큰 수준의 입력 길이 제한을 가지므로 긴 특허문서를 요약이나 절단 없이 벡터로 표현하는 데 한계가 있다. 이에 본 연구는 장문 입력이 가능한 다국어 임베딩 모델을 한국어 특허 데이터로 미세조정하여, 특허 분야 검색에 활용 가능한 임베딩 모델을 제시한다. 임베딩 모델 미세조정을 위해 AI Hub에서 제공하는 약 127만 건의 대규모 한국어 특허 데이터셋을 활용하였으며, 최대 8,192 토큰을 지원하는 BGE-M3 모델을 기반으로 미세조정을 수행하였다. 효과적인 도메인 지식 학습을 위해 국제특허분류(International Patent Classification, IPC) 체계와 키워드 일치도를 기준으로 학습 샘플의 상대적 난이도를 5단계로 세분화한 커리큘럼 러닝(curriculum learning) 전략을 채택하였다. 또한 절대적 목표 거리 제약조건을 추가한 트리플렛 손실함수(triplet loss)를 결합하여, 모델이 특허 간의 미세한 기술적 차이와 연속적인 유사도 스펙트럼을 정교하게 학습하도록 유도하였다. 더불어, 특정 도메인으로의 미세조정 시 흔히 발생하는 범용 성능의 저하 문제를 완화하기 위해 사전 훈련된 한국어 범용 임베딩 모델과의 사후 보간(post-hoc interpolation)을 수행하였다. 성능평가 결과, 보간 계수 α=0.3을 적용한 모델이 특허 도메인 내 정밀도(Precision@k) 및 인용 특허 변별력에서 높은 성능을 보였을 뿐만 아니라, 범용 한국어 검색 벤치마크에서도 기존 범용 다국어 모델들과 유사한 수준의 경쟁력을 유지함을 확인하였다. 추가적으로 정성평가 사례를 통해 512 토큰 이후에 위치한 핵심 기술요소가 검색 결과에 영향을 미칠 수 있음을 확인하였다. 나아가 본 연구는 제안된 임베딩 모델을 검색증강생성 기반 특허 질의응답 시스템에 통합하고 RAGAs(Retrieval Augmented Generation Assessment) 프레임워크를 활용하여 질의응답 성능을 평가하였다. 또한 벡터 데이터베이스 기반 검색 모듈(retriever), 경량 리랭커(reranker), 답변 생성 에이전트, 채팅 인터페이스를 결합한 아키텍처를 구현함으로써, 신규성 평가 및 유사 특허정보 추출과 같은 실무 작업에 응용할 수 있는 예시를 제공한다는 점에서 실용적 의의를 가진다.
    번역하기

    최근 전 세계적인 기술 경쟁 심화로 특허 출원이 지속적으로 증가함에 따라, 방대한 특허 문헌 속에서 필요한 정보를 효과적으로 탐색하고 분석하는 기술의 중요성이 대두되고 있다. 그러나...

    최근 전 세계적인 기술 경쟁 심화로 특허 출원이 지속적으로 증가함에 따라, 방대한 특허 문헌 속에서 필요한 정보를 효과적으로 탐색하고 분석하는 기술의 중요성이 대두되고 있다. 그러나 특허 분야에서 공개된 기존 언어모델은 주로 특허 말뭉치로 사전학습된 모델이며, 문서 간 유사도 검색을 위해 임베딩 공간을 최적화한 공개 모델은 부족한 실정이다. 또한 대부분 512 토큰 수준의 입력 길이 제한을 가지므로 긴 특허문서를 요약이나 절단 없이 벡터로 표현하는 데 한계가 있다. 이에 본 연구는 장문 입력이 가능한 다국어 임베딩 모델을 한국어 특허 데이터로 미세조정하여, 특허 분야 검색에 활용 가능한 임베딩 모델을 제시한다. 임베딩 모델 미세조정을 위해 AI Hub에서 제공하는 약 127만 건의 대규모 한국어 특허 데이터셋을 활용하였으며, 최대 8,192 토큰을 지원하는 BGE-M3 모델을 기반으로 미세조정을 수행하였다. 효과적인 도메인 지식 학습을 위해 국제특허분류(International Patent Classification, IPC) 체계와 키워드 일치도를 기준으로 학습 샘플의 상대적 난이도를 5단계로 세분화한 커리큘럼 러닝(curriculum learning) 전략을 채택하였다. 또한 절대적 목표 거리 제약조건을 추가한 트리플렛 손실함수(triplet loss)를 결합하여, 모델이 특허 간의 미세한 기술적 차이와 연속적인 유사도 스펙트럼을 정교하게 학습하도록 유도하였다. 더불어, 특정 도메인으로의 미세조정 시 흔히 발생하는 범용 성능의 저하 문제를 완화하기 위해 사전 훈련된 한국어 범용 임베딩 모델과의 사후 보간(post-hoc interpolation)을 수행하였다. 성능평가 결과, 보간 계수 α=0.3을 적용한 모델이 특허 도메인 내 정밀도(Precision@k) 및 인용 특허 변별력에서 높은 성능을 보였을 뿐만 아니라, 범용 한국어 검색 벤치마크에서도 기존 범용 다국어 모델들과 유사한 수준의 경쟁력을 유지함을 확인하였다. 추가적으로 정성평가 사례를 통해 512 토큰 이후에 위치한 핵심 기술요소가 검색 결과에 영향을 미칠 수 있음을 확인하였다. 나아가 본 연구는 제안된 임베딩 모델을 검색증강생성 기반 특허 질의응답 시스템에 통합하고 RAGAs(Retrieval Augmented Generation Assessment) 프레임워크를 활용하여 질의응답 성능을 평가하였다. 또한 벡터 데이터베이스 기반 검색 모듈(retriever), 경량 리랭커(reranker), 답변 생성 에이전트, 채팅 인터페이스를 결합한 아키텍처를 구현함으로써, 신규성 평가 및 유사 특허정보 추출과 같은 실무 작업에 응용할 수 있는 예시를 제공한다는 점에서 실용적 의의를 가진다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    As global technological competition intensifies and patent filings continue to increase, the ability to efficiently explore and analyze relevant information within large-scale patent corpora has become increasingly important. However, existing publicly available language models in the patent domain are mainly models pretrained on patent corpora, while models whose embedding spaces are optimized for document-level similarity search remain limited. In addition, most of these models are constrained by input length limits of around 512 tokens, making it difficult to represent long patent documents as vectors without summarization or truncation. To address these limitations, this study fine-tunes a long-context multilingual embedding model using Korean patent data and presents an embedding model for Korean patent search. For fine-tuning, this study utilizes a large-scale Korean patent dataset of approximately 1.27 million documents provided by AI Hub and adopts BGE-M3, which supports input lengths of up to 8,192 tokens, as the base model. To facilitate effective domain-specific representation learning, this study introduces a curriculum learning strategy that stratifies the relative difficulty of training samples into five levels based on the International Patent Classification (IPC) scheme and keyword matching. In addition, triplet loss is combined with an absolute target-distance constraint, encouraging the model to capture fine-grained technical differences between patents and reflect similarity as a continuous spectrum. Furthermore, to mitigate the degradation of general-purpose performance that commonly occurs during domain-specific fine-tuning, this study performs post-hoc weight interpolation with a pretrained Korean general-purpose embedding model. Experimental results show that the model with an interpolation coefficient of α=0.3 achieves competitive performance in patent-domain precision (Precision@k) and cited-patent discrimination, while maintaining performance comparable to that of existing general-purpose multilingual models on Korean retrieval benchmarks. Additionally, qualitative evaluation cases show that key technical elements located beyond the first 512 tokens can affect retrieval results. Finally, the proposed embedding model is integrated into a retrieval-augmented generation (RAG)-based patent question answering system, and its question answering performance is evaluated using the Retrieval Augmented Generation Assessment (RAGAs) framework. By implementing an architecture that combines vector database retrieval, a lightweight large language model (LLM) reranker, an answer-generation agent, and a chat interface, this study provides a practical example that can be applied to real-world tasks such as novelty assessment and extraction of similar patent information.
    번역하기

    As global technological competition intensifies and patent filings continue to increase, the ability to efficiently explore and analyze relevant information within large-scale patent corpora has become increasingly important. However, existing publicl...

    As global technological competition intensifies and patent filings continue to increase, the ability to efficiently explore and analyze relevant information within large-scale patent corpora has become increasingly important. However, existing publicly available language models in the patent domain are mainly models pretrained on patent corpora, while models whose embedding spaces are optimized for document-level similarity search remain limited. In addition, most of these models are constrained by input length limits of around 512 tokens, making it difficult to represent long patent documents as vectors without summarization or truncation. To address these limitations, this study fine-tunes a long-context multilingual embedding model using Korean patent data and presents an embedding model for Korean patent search. For fine-tuning, this study utilizes a large-scale Korean patent dataset of approximately 1.27 million documents provided by AI Hub and adopts BGE-M3, which supports input lengths of up to 8,192 tokens, as the base model. To facilitate effective domain-specific representation learning, this study introduces a curriculum learning strategy that stratifies the relative difficulty of training samples into five levels based on the International Patent Classification (IPC) scheme and keyword matching. In addition, triplet loss is combined with an absolute target-distance constraint, encouraging the model to capture fine-grained technical differences between patents and reflect similarity as a continuous spectrum. Furthermore, to mitigate the degradation of general-purpose performance that commonly occurs during domain-specific fine-tuning, this study performs post-hoc weight interpolation with a pretrained Korean general-purpose embedding model. Experimental results show that the model with an interpolation coefficient of α=0.3 achieves competitive performance in patent-domain precision (Precision@k) and cited-patent discrimination, while maintaining performance comparable to that of existing general-purpose multilingual models on Korean retrieval benchmarks. Additionally, qualitative evaluation cases show that key technical elements located beyond the first 512 tokens can affect retrieval results. Finally, the proposed embedding model is integrated into a retrieval-augmented generation (RAG)-based patent question answering system, and its question answering performance is evaluated using the Retrieval Augmented Generation Assessment (RAGAs) framework. By implementing an architecture that combines vector database retrieval, a lightweight large language model (LLM) reranker, an answer-generation agent, and a chat interface, this study provides a practical example that can be applied to real-world tasks such as novelty assessment and extraction of similar patent information.

    더보기

    목차 (Table of Contents)

    • 표 목차 iv
    • 그림 목차 v
    • 국문요지 vi
    • 제 1 장 서론 1
    • 표 목차 iv
    • 그림 목차 v
    • 국문요지 vi
    • 제 1 장 서론 1
    • 1.1. 연구배경 1
    • 1.2. 연구목적 및 범위 4
    • 1.3. 연구 방법 10
    • 제 2 장 이론적 배경 및 선행연구 16
    • 2.1. 특허 검색 17
    • 2.1.1. 특허 검색 방법 17
    • 2.1.2. 메타데이터 기반 검색 19
    • 2.1.3. 키워드 기반 검색 21
    • 2.1.4. 의미 기반 검색 25
    • 2.1.5. 특허 검색 방법의 발전 27
    • 2.2. 언어모델 29
    • 2.3. 특허 분야 언어모델 34
    • 2.4. 효율적 미세조정 기법 44
    • 2.5. 미세조정의 손실함수 46
    • 2.6. 커리큘럼 러닝 50
    • 2.7. 검색증강생성(RAG)과 임베딩 모델 54
    • 2.8. Agentic AI 58
    • 제 3 장 학습 데이터 및 훈련 전략 65
    • 3.1. 학습 데이터 66
    • 3.2. 임베딩 모델 69
    • 3.3. 훈련 전략 72
    • 3.4. 사후 보간 (post-hoc interpolation) 81
    • 제 4 장 임베딩 모델 성능평가 84
    • 4.1. 성능평가 지표 (특허 영역) 84
    • 4.2. 성능평가 지표 (범용) 88
    • 4.3. 평가 결과 89
    • 4.4. 정성평가 99
    • 제 5 장 질의응답 평가 112
    • 5.1. 평가 방법 112
    • 5.2. 질의 생성 및 응답 시나리오 114
    • 5.3. 성능평가 지표 116
    • 5.4. 평가 결과 119
    • 제 6 장 RAG 기반 질의응답 시스템 구성 123
    • 6.1. 아키텍처 123
    • 6.2. 상세 구성 예시 126
    • 제 7 장 결론 131
    • 7.1. 연구 요약 131
    • 7.2. 연구의 의의 134
    • 7.3. 한계점 및 향후 연구 방향 137
    • 참고문헌 142
    • 부록 163
    • ABSTRACT 164
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼