RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    조선시대 한문 사료를 위한의미 기반 검색 모델의 개발과 활용―『조선왕조실록』과 『경국대전』을 중심으로― = Development and Application of a Semantic Search Model for Classical Chinese Sources of the Joseon Dynasty : Focusing on the Veritable Records of the Joseon Dynasty and the Grand Canon of State Administration

    한글로보기

    https://www.riss.kr/link?id=A110219032

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    조선시대 사료는 대부분 한문으로 이루어져 있으며, 그중 󰡔조선왕조실록󰡕은 조선 전시대의 다양한 내용을 담고 있어 역사 연구의 기본이 된다. 또한 󰡔경국대전󰡕은 조선 초기 제도를 집대성하고 통치 이념을 담고 있는 서적인데, 조문이 간략하고 함축되어 있어 쉽게 이해하기 어렵다. 이를 쉽게 이해하기 위해 󰡔경국대전󰡕 조문의 입법 의도를 실록 속에서 찾아 검토할 필요가 있다. 그러나 󰡔경국대전󰡕 조문의 내용은 실록에서 같은 표현으로 나타나지 않는 경우가 많아 현재 실록 웹사이트의 키워드 검색만으로는 안정적으로 수행하기 어렵다.
    따라서 󰡔경국대전󰡕 조문을 질의(query)로 입력하면 문자열의 일치 여부와 상관없이 관련 실록 기사를 검색할 수 있는 의미 기반 검색 모델을 개발하였다. 사고전서를 학습한 SikuRoBERTa 모델을 베이스로 하여 실록과 󰡔승정원일기󰡕로 MLM 학습 후 조선시대 문체와 어휘를 반영하도록 하였다. 이어 비지도 SimCSE와 약지도 대조학습을 통해 문장/청크 임베딩을 가능하게 하였다. 검색 단계에서는 의미 유사도(dense) 기반 결과와 n-gram TF-IDF(sparse) 기반 결과를 RRF로 결합하고, 상위 결과를 seed로 삼는 2단계 재검색(2-hop)을 도입하여 누락되는 기사가 적게 하였다. 실험 결과 Capped Recall@10 0.716, HitRate@10 0.9의 성능을 보였다. 이와 같이 AI 인문학 도구로서 의미 기반 검색 모델을 제안한다.
    번역하기

    조선시대 사료는 대부분 한문으로 이루어져 있으며, 그중 󰡔조선왕조실록󰡕은 조선 전시대의 다양한 내용을 담고 있어 역사 연구의 기본이 된다. 또한 󰡔경국대전󰡕은 조선 초기 제도�...

    조선시대 사료는 대부분 한문으로 이루어져 있으며, 그중 󰡔조선왕조실록󰡕은 조선 전시대의 다양한 내용을 담고 있어 역사 연구의 기본이 된다. 또한 󰡔경국대전󰡕은 조선 초기 제도를 집대성하고 통치 이념을 담고 있는 서적인데, 조문이 간략하고 함축되어 있어 쉽게 이해하기 어렵다. 이를 쉽게 이해하기 위해 󰡔경국대전󰡕 조문의 입법 의도를 실록 속에서 찾아 검토할 필요가 있다. 그러나 󰡔경국대전󰡕 조문의 내용은 실록에서 같은 표현으로 나타나지 않는 경우가 많아 현재 실록 웹사이트의 키워드 검색만으로는 안정적으로 수행하기 어렵다.
    따라서 󰡔경국대전󰡕 조문을 질의(query)로 입력하면 문자열의 일치 여부와 상관없이 관련 실록 기사를 검색할 수 있는 의미 기반 검색 모델을 개발하였다. 사고전서를 학습한 SikuRoBERTa 모델을 베이스로 하여 실록과 󰡔승정원일기󰡕로 MLM 학습 후 조선시대 문체와 어휘를 반영하도록 하였다. 이어 비지도 SimCSE와 약지도 대조학습을 통해 문장/청크 임베딩을 가능하게 하였다. 검색 단계에서는 의미 유사도(dense) 기반 결과와 n-gram TF-IDF(sparse) 기반 결과를 RRF로 결합하고, 상위 결과를 seed로 삼는 2단계 재검색(2-hop)을 도입하여 누락되는 기사가 적게 하였다. 실험 결과 Capped Recall@10 0.716, HitRate@10 0.9의 성능을 보였다. 이와 같이 AI 인문학 도구로서 의미 기반 검색 모델을 제안한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Most historical sources from the Joseon period are written in Classical Chinese, and among them, the Veritable Records of the Joseon Dynasty serves as the foundation for historical research, as it contains diverse records spanning the entire dynasty. The Grand Canon of Administrating the State is likewise a book that compiles the early Joseon institutions and articulates its governing ideology; however, its articles are concise and highly condensed, which makes them difficult to understand. In order to facilitate interpretation, it is necessary to trace and examine the legislative intent of the Grand Canon of Administrating the State within the Veritable Records of the Joseon Dynasty. Yet the content of these articles often does not appear in the Veritable Records of the Joseon Dynasty in the same wording, and for this reason keyword-based search on existing web services cannot support such work in a stable manner.
    Accordingly, this study develops a semantic search model that retrieves relevant articles from the Veritable Records of the Joseon Dynasty even when there is no direct string-level match, using the articles of the Grand Canon of Administrating the State as queries. Building on SikuRoBERTa pretrained on the Siku Quanshu, the model is further adapted via masked language modeling (MLM) on the Veritable Records of the Joseon Dynasty and the Daily Records of Royal Secretariat of Joseon Dynasty so as to reflect Joseon-period style and vocabulary. The model then enables sentence/chunk embedding through unsupervised SimCSE and weakly supervised contrastive learning. In the retrieval stage, dense semantic-similarity results and sparse n-gram TF–IDF results are fused with reciprocal rank fusion (RRF), and a two-hop re-retrieval scheme is introduced that uses top-ranked results as seeds to reduce omissions. Experiments yield Capped Recall@10 of 0.716 and HitRate@10 of 0.9. In this way, the study proposes a semantic search model as a tool for AI Humanities.
    번역하기

    Most historical sources from the Joseon period are written in Classical Chinese, and among them, the Veritable Records of the Joseon Dynasty serves as the foundation for historical research, as it contains diverse records spanning the entire dynasty. ...

    Most historical sources from the Joseon period are written in Classical Chinese, and among them, the Veritable Records of the Joseon Dynasty serves as the foundation for historical research, as it contains diverse records spanning the entire dynasty. The Grand Canon of Administrating the State is likewise a book that compiles the early Joseon institutions and articulates its governing ideology; however, its articles are concise and highly condensed, which makes them difficult to understand. In order to facilitate interpretation, it is necessary to trace and examine the legislative intent of the Grand Canon of Administrating the State within the Veritable Records of the Joseon Dynasty. Yet the content of these articles often does not appear in the Veritable Records of the Joseon Dynasty in the same wording, and for this reason keyword-based search on existing web services cannot support such work in a stable manner.
    Accordingly, this study develops a semantic search model that retrieves relevant articles from the Veritable Records of the Joseon Dynasty even when there is no direct string-level match, using the articles of the Grand Canon of Administrating the State as queries. Building on SikuRoBERTa pretrained on the Siku Quanshu, the model is further adapted via masked language modeling (MLM) on the Veritable Records of the Joseon Dynasty and the Daily Records of Royal Secretariat of Joseon Dynasty so as to reflect Joseon-period style and vocabulary. The model then enables sentence/chunk embedding through unsupervised SimCSE and weakly supervised contrastive learning. In the retrieval stage, dense semantic-similarity results and sparse n-gram TF–IDF results are fused with reciprocal rank fusion (RRF), and a two-hop re-retrieval scheme is introduced that uses top-ranked results as seeds to reduce omissions. Experiments yield Capped Recall@10 of 0.716 and HitRate@10 of 0.9. In this way, the study proposes a semantic search model as a tool for AI Humanities.

    더보기

    동일학술지(권/호) 다른 논문

    동일학술지 더보기

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼