RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    공간-텍스트 지식그래프 구축을 통한 GraphRAG 기반 공간 질의응답 연구 = GraphRAG-based Geospatial Question Answering via Construction of Geo-Textual Knowledge Graph

    한글로보기

    https://www.riss.kr/link?id=T17451235

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    공간 질의응답은 지리적 개체와 공간 관계를 함께 이해해야 하기 때문에, 일반적인 검색 증강 생성(RAG) 기반의 거대 언어 모델은 구조적 제약을 처리하는 데 한계를 가진다. 본 연구는 이러한 한계를 극복하기 위해 공간지식그래프와 한글 Wikipedia 문서를 결합한 공간-텍스트 지식그래프(Geo-Textual Knowledge Graph, GTKG)를 구축하고, 이를 활용한 그래프 기반 검색 증강 생성(GraphRAG) 파이프라인을 제안한다.
    제안하는 GraphRAG 파이프라인은 그래프 탐색과 의미 기반 문서 검색을 단계적으로 결합한 계층적 검색 구조를 가진다. 우선 그래프 탐색을 통해 공간적 제약 조건을 만족하는 후보 개체를 한정하고, 이후 해당 개체에 연결된 문서 청크를 의미적으로 순위화한다. 특히, 그래프 탐색 결과로 다수의 개체가 반환되는 복합 요약 질의의 경우, 특정 개체로의 정보 편중을 방지하고 균형 잡힌 근거를 확보하고자 라운드 로빈 병합 방식을 채택하였다.
    실험은 서울특별시 서초구를 대상으로 수행하였다. 브이월드(VWorld) 데이터와 OpenStreetMap(OSM) 데이터를 이용하여 공간지식그래프를 구성하고, 이를 서울시 전역의 한글 Wikipedia 문서와 결합하였다. 평가 데이터셋은 질의 유형에 따라 그래프 기반, 문서 기반, 단일 개체 결합, 다중 개체 결합의 네 가지 범주로 구분하였다. 각 질의에 대해 NaiveRAG와 GraphRAG의 성능을 비교하였다.
    평가 결과, GraphRAG는 모든 질의 유형에서 NaiveRAG보다 높은 검색 및 답변 품질을 보였다. 특히 검색 재현율(Recall@10)을 기준으로, 복합 질의에 해당하는 단일 개체 결합 질의에서 96.9%, 다중 개체 결합 질의에서 89.2%를 기록하였다. 이러한 검색 성능의 향상은 EM, F1, ROUGE-L 등의 답변 품질 지표에서도 유의한 개선으로 이어졌다. 이는 GraphRAG의 그래프 탐색 단계가 불필요한 문서 검색 범위를 줄였을 뿐만 아니라, 명확한 사전 맥락을 제공하여 LLM이 동일한 정보도 더 정확하게 해석하도록 유도한 효과로 분석된다.
    본 연구는 제안한 GraphRAG 파이프라인을 공간 도메인에 적용하여, 구조적 관계 추론이 필수적인 복합 질의에 대한 거대 언어 모델의 한계를 극복할 수 있음을 보였다. 나아가 공공 데이터를 활용하여 실제 행정구 단위의 GTKG를 구현하고 그 성능을 입증함으로써, 본 방법론의 실용적 가치와 확장 가능성을 실증하였다.
    향후 연구로는 자연어 질의의 Cypher 자동 변환 기능 통합과 더불어, 본 연구에서 확인된 의미론적 검색의 한계를 보완하기 위한 하이브리드 검색을 도입하는 것을 후속 과제로 제시한다.
    번역하기

    공간 질의응답은 지리적 개체와 공간 관계를 함께 이해해야 하기 때문에, 일반적인 검색 증강 생성(RAG) 기반의 거대 언어 모델은 구조적 제약을 처리하는 데 한계를 가진다. 본 연구는 이러...

    공간 질의응답은 지리적 개체와 공간 관계를 함께 이해해야 하기 때문에, 일반적인 검색 증강 생성(RAG) 기반의 거대 언어 모델은 구조적 제약을 처리하는 데 한계를 가진다. 본 연구는 이러한 한계를 극복하기 위해 공간지식그래프와 한글 Wikipedia 문서를 결합한 공간-텍스트 지식그래프(Geo-Textual Knowledge Graph, GTKG)를 구축하고, 이를 활용한 그래프 기반 검색 증강 생성(GraphRAG) 파이프라인을 제안한다.
    제안하는 GraphRAG 파이프라인은 그래프 탐색과 의미 기반 문서 검색을 단계적으로 결합한 계층적 검색 구조를 가진다. 우선 그래프 탐색을 통해 공간적 제약 조건을 만족하는 후보 개체를 한정하고, 이후 해당 개체에 연결된 문서 청크를 의미적으로 순위화한다. 특히, 그래프 탐색 결과로 다수의 개체가 반환되는 복합 요약 질의의 경우, 특정 개체로의 정보 편중을 방지하고 균형 잡힌 근거를 확보하고자 라운드 로빈 병합 방식을 채택하였다.
    실험은 서울특별시 서초구를 대상으로 수행하였다. 브이월드(VWorld) 데이터와 OpenStreetMap(OSM) 데이터를 이용하여 공간지식그래프를 구성하고, 이를 서울시 전역의 한글 Wikipedia 문서와 결합하였다. 평가 데이터셋은 질의 유형에 따라 그래프 기반, 문서 기반, 단일 개체 결합, 다중 개체 결합의 네 가지 범주로 구분하였다. 각 질의에 대해 NaiveRAG와 GraphRAG의 성능을 비교하였다.
    평가 결과, GraphRAG는 모든 질의 유형에서 NaiveRAG보다 높은 검색 및 답변 품질을 보였다. 특히 검색 재현율(Recall@10)을 기준으로, 복합 질의에 해당하는 단일 개체 결합 질의에서 96.9%, 다중 개체 결합 질의에서 89.2%를 기록하였다. 이러한 검색 성능의 향상은 EM, F1, ROUGE-L 등의 답변 품질 지표에서도 유의한 개선으로 이어졌다. 이는 GraphRAG의 그래프 탐색 단계가 불필요한 문서 검색 범위를 줄였을 뿐만 아니라, 명확한 사전 맥락을 제공하여 LLM이 동일한 정보도 더 정확하게 해석하도록 유도한 효과로 분석된다.
    본 연구는 제안한 GraphRAG 파이프라인을 공간 도메인에 적용하여, 구조적 관계 추론이 필수적인 복합 질의에 대한 거대 언어 모델의 한계를 극복할 수 있음을 보였다. 나아가 공공 데이터를 활용하여 실제 행정구 단위의 GTKG를 구현하고 그 성능을 입증함으로써, 본 방법론의 실용적 가치와 확장 가능성을 실증하였다.
    향후 연구로는 자연어 질의의 Cypher 자동 변환 기능 통합과 더불어, 본 연구에서 확인된 의미론적 검색의 한계를 보완하기 위한 하이브리드 검색을 도입하는 것을 후속 과제로 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Geospatial question answering requires an understanding of both geographic entities and their spatial relationships, which poses challenges for conventional Retrieval-Augmented Generation(RAG) models built upon large language models(LLMs). To address this limitation, this study proposes a Graph-based Retrieval-Augmented Generation(GraphRAG) pipeline that integrates a geographic knowledge graph with Korean Wikipedia documents.
    The proposed GraphRAG pipeline adopts a hierarchical retrieval framework that sequentially combines graph traversal and semantic document ranking. The system first filters candidate entities that satisfy spatial constraints through graph traversal, and then ranks the document chunks connected to those entities based on semantic relevance. Notably, for complex summary queries where graph traversal returns multiple entities, a round-robin merging strategy was adopted to prevent information skew toward specific entities and to ensure balanced evidence.
    The experiment was conducted in Seocho-gu, Seoul. A geographic knowledge graph was constructed using data from VWorld and OpenStreetMap and was linked with Korean Wikipedia documents covering the entire Seoul metropolitan area. An evaluation dataset was custom-built and categorized into four types based on the required information sources: graph-based, document-based, single-entity integrated, and multi-entity integrated. The performance of GraphRAG was compared with a baseline NaiveRAG model across these query types.
    Experimental results demonstrate that GraphRAG consistently outperforms NaiveRAG in both retrieval and answer quality across all query types. In particular, based on Recall@10, it achieved 96.9% for single-entity integrated queries, a form of complex query, and 89.2% for multi-entity integrated queries. These retrieval improvements also translated into significant enhancements in answer quality metrics such as EM, F1, and ROUGE-L. This is attributed not only to the graph traversal stage reducing the irrelevant search space but also to its effect of providing clear prior context, which guides the LLM to interpret the same information more accurately.
    This study demonstrated that by applying the proposed GraphRAG pipeline to the geospatial domain, it is possible to overcome the limitations of large language models for complex queries requiring structural relation reasoning. Furthermore, by implementing a city-scale GTKG(Geo-Textual Knowledge Graph) using public data and verifying its performance, this study substantiates the practical value and scalability of the proposed methodology.
    Future work suggests integrating a natural language to Cypher automatic conversion function, as well as introducing Hybrid Search to compensate for the limitations of semantic search identified in this study.
    번역하기

    Geospatial question answering requires an understanding of both geographic entities and their spatial relationships, which poses challenges for conventional Retrieval-Augmented Generation(RAG) models built upon large language models(LLMs). To address ...

    Geospatial question answering requires an understanding of both geographic entities and their spatial relationships, which poses challenges for conventional Retrieval-Augmented Generation(RAG) models built upon large language models(LLMs). To address this limitation, this study proposes a Graph-based Retrieval-Augmented Generation(GraphRAG) pipeline that integrates a geographic knowledge graph with Korean Wikipedia documents.
    The proposed GraphRAG pipeline adopts a hierarchical retrieval framework that sequentially combines graph traversal and semantic document ranking. The system first filters candidate entities that satisfy spatial constraints through graph traversal, and then ranks the document chunks connected to those entities based on semantic relevance. Notably, for complex summary queries where graph traversal returns multiple entities, a round-robin merging strategy was adopted to prevent information skew toward specific entities and to ensure balanced evidence.
    The experiment was conducted in Seocho-gu, Seoul. A geographic knowledge graph was constructed using data from VWorld and OpenStreetMap and was linked with Korean Wikipedia documents covering the entire Seoul metropolitan area. An evaluation dataset was custom-built and categorized into four types based on the required information sources: graph-based, document-based, single-entity integrated, and multi-entity integrated. The performance of GraphRAG was compared with a baseline NaiveRAG model across these query types.
    Experimental results demonstrate that GraphRAG consistently outperforms NaiveRAG in both retrieval and answer quality across all query types. In particular, based on Recall@10, it achieved 96.9% for single-entity integrated queries, a form of complex query, and 89.2% for multi-entity integrated queries. These retrieval improvements also translated into significant enhancements in answer quality metrics such as EM, F1, and ROUGE-L. This is attributed not only to the graph traversal stage reducing the irrelevant search space but also to its effect of providing clear prior context, which guides the LLM to interpret the same information more accurately.
    This study demonstrated that by applying the proposed GraphRAG pipeline to the geospatial domain, it is possible to overcome the limitations of large language models for complex queries requiring structural relation reasoning. Furthermore, by implementing a city-scale GTKG(Geo-Textual Knowledge Graph) using public data and verifying its performance, this study substantiates the practical value and scalability of the proposed methodology.
    Future work suggests integrating a natural language to Cypher automatic conversion function, as well as introducing Hybrid Search to compensate for the limitations of semantic search identified in this study.

    더보기

    목차 (Table of Contents)

    • 1. 서 론 1
    • 1.1 연구의 배경 및 필요성 1
    • 1.2 관련 연구의 한계 2
    • 1.2.1 NaiveRAG의 한계 2
    • 1.2.2 KBQA의 한계 2
    • 1. 서 론 1
    • 1.1 연구의 배경 및 필요성 1
    • 1.2 관련 연구의 한계 2
    • 1.2.1 NaiveRAG의 한계 2
    • 1.2.2 KBQA의 한계 2
    • 1.2.3 공간 QA 연구의 동향 및 한계 3
    • 1.3 연구 목적 및 질문 6
    • 1.4 연구의 범위 및 구성 7
    • 2. 이론적 배경 8
    • 2.1 공간 질의응답 연구 동향 9
    • 2.1.1 GeoKBQA 9
    • 2.1.2 Spatial-RAG 10
    • 2.1.3 GeoGraphRAG 11
    • 2.2 HGeoKG와 TAG 기반의 지식 모델링 13
    • 2.2.1 HGeoKG 13
    • 2.2.2 텍스트 속성 그래프(TAG) 14
    • 2.3 합성 QA 데이터셋 연구 16
    • 3. 방 법 론 18
    • 3.1 데이터 수집 및 전처리 20
    • 3.1.1 브이월드 행정구역 데이터 21
    • 3.1.2 OpenStreetMap 데이터 22
    • 3.1.3 한글 Wikipedia 데이터 24
    • 3.2 서초구 공간지식그래프 구축 26
    • 3.3 서초구 공간-텍스트 지식그래프 구축 30
    • 3.3.1 문서 청킹 30
    • 3.3.2 노드 및 속성 매핑 31
    • 3.3.3 그래프 통합 31
    • 3.4 GraphRAG 구성 요소 구현 34
    • 3.4.1 G-Indexing 35
    • 3.4.2 G-Retrieval 35
    • 3.4.3 G-Generation 39
    • 4. 실험 및 평가 41
    • 4.1 평가용 QA 데이터셋 구축 42
    • 4.1.1 Graph-only QA 데이터셋 생성 44
    • 4.1.2 Document-only QA 데이터셋 생성 46
    • 4.1.3 Integrated QA(Single entity) 데이터셋 생성 47
    • 4.1.4 Integrated QA(Multi entity) 데이터셋 생성 48
    • 4.2 평가지표 설정 51
    • 4.2.1 답변 품질 평가지표 51
    • 4.2.2 검색 성능 평가지표 53
    • 4.2.3 효율성 평가지표 54
    • 4.3 실험 결과 분석 56
    • 4.3.1 Type 1(Graph-only QA) 성능 분석 56
    • 4.3.2 Type 2(Document-only QA) 성능 분석 57
    • 4.3.3 Type 3(Integrated QA - Single entity) 성능 분석 59
    • 4.3.4 Type 4(Integrated QA - Multi entity) 성능 분석 62
    • 4.3.5 효율성 평가 64
    • 4.4 종합 논의 및 시사점 66
    • 5. 결 론 68
    • A. 부 록 70
    • A.1 검색 전략 비교를 위한 LLM-as-a-Judge 평가 70
    • A.2 실험에 사용된 프롬프트 원문 72
    • A.2.1 QA 데이터셋 생성 프롬프트 72
    • A.2.2 답변 생성 프롬프트 73
    • A.3 평가 데이터셋 전체 목록 75
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼