RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    지식 그래프 추론과 전문 문헌 근거 검색을 활용한 바이오메디컬 QA 벤치마크 자동 생성 = Bio-KGR: Automating Biomedical QA Benchmarks with Knowledge Graph Reasoning and Evidence Retrieva

    한글로보기

    https://www.riss.kr/link?id=T17380612

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 대규모 언어 모델(Large Language Model; LLM)의 발전과 함께 바이오메디컬 분야 연구가 가속화되면서 고품질 질문-답변(QA) 데이터셋에 대한 수요가 빠르게 늘고 있다. 그러나 기존 벤치마크는 수작업 질문과 고정된 지식 그래프에 의존해 최신성·확장성·효율성에 한계를 보인다. 특히, 이 분야는 계층적 추론 능력과 전문 지식을 요구하며 전문 문헌 기반의 근거가 필요하기 때문에 표면적 지식 질문 만으로 LLM의 추론 능력을 정확히 평가하기 어렵다. 이러한 한계를 해결하기 위해 Bio-KGR 프레임워크를 제안한다. 이 프레임워크는 도메인 특화 지식 그래프와 검색 증강 생성을 통해 PubMed 문헌에서 검색된 근거를 활용하여 다양한 유형의 개방형 QA 쌍을 자동으로 생성한다. 소규모 그래프에도 적용 가능하며, 문헌 기반 필터링과 난이도 조절을 통해 데이터 품질을 확보한다. 생성된 QA 쌍은 다양한 LLM 및 전문가 평가를 통하여 질문 자연스러움, 답변 적절성, 추론 난이도, 데이터셋 다양성이라는 네 가지 기준으로 분석되었다. 실험 결과, Bio-KGR는 제로/퓨샷 설정 및 기존 수작업 벤치마크와 비교했을 때도 대부분의 지표에서 우수한 성능을 보였다. 또한 모델 간 일관된 평가 결과와 다양한 난이도 구성으로 Bio-KGR는 자동 생성임에도 불구하고 높은 신뢰성과 품질을 입증했다. 이러한 결과는 LLM의 고차원적 추론 능력을 체계적으로 평가할 수 있는 바이오메디컬 QA 벤치마크를 수작업 없이 자동으로 구축할 수 있음을 보여준다.
    번역하기

    최근 대규모 언어 모델(Large Language Model; LLM)의 발전과 함께 바이오메디컬 분야 연구가 가속화되면서 고품질 질문-답변(QA) 데이터셋에 대한 수요가 빠르게 늘고 있다. 그러나 기존 벤치마크는...

    최근 대규모 언어 모델(Large Language Model; LLM)의 발전과 함께 바이오메디컬 분야 연구가 가속화되면서 고품질 질문-답변(QA) 데이터셋에 대한 수요가 빠르게 늘고 있다. 그러나 기존 벤치마크는 수작업 질문과 고정된 지식 그래프에 의존해 최신성·확장성·효율성에 한계를 보인다. 특히, 이 분야는 계층적 추론 능력과 전문 지식을 요구하며 전문 문헌 기반의 근거가 필요하기 때문에 표면적 지식 질문 만으로 LLM의 추론 능력을 정확히 평가하기 어렵다. 이러한 한계를 해결하기 위해 Bio-KGR 프레임워크를 제안한다. 이 프레임워크는 도메인 특화 지식 그래프와 검색 증강 생성을 통해 PubMed 문헌에서 검색된 근거를 활용하여 다양한 유형의 개방형 QA 쌍을 자동으로 생성한다. 소규모 그래프에도 적용 가능하며, 문헌 기반 필터링과 난이도 조절을 통해 데이터 품질을 확보한다. 생성된 QA 쌍은 다양한 LLM 및 전문가 평가를 통하여 질문 자연스러움, 답변 적절성, 추론 난이도, 데이터셋 다양성이라는 네 가지 기준으로 분석되었다. 실험 결과, Bio-KGR는 제로/퓨샷 설정 및 기존 수작업 벤치마크와 비교했을 때도 대부분의 지표에서 우수한 성능을 보였다. 또한 모델 간 일관된 평가 결과와 다양한 난이도 구성으로 Bio-KGR는 자동 생성임에도 불구하고 높은 신뢰성과 품질을 입증했다. 이러한 결과는 LLM의 고차원적 추론 능력을 체계적으로 평가할 수 있는 바이오메디컬 QA 벤치마크를 수작업 없이 자동으로 구축할 수 있음을 보여준다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recently, with the advancement of large language models (LLMs), biomedical research has been accelerating, leading to a rapidly increasing demand for high-quality question-answering (QA) datasets. However, existing biomedical QA benchmarks rely on manually created questions and fixed large-scale knowledge graphs, showing limitations in terms of currency, scalability, and efficiency. Since the biomedical field requires hierarchical reasoning capabilities and high-level domain expertise, with evidence based on specialized literature being essential, it is difficult to accurately evaluate LLMs' reasoning abilities using only superficial knowledge questions. To address these issues, we propose the Bio-KGR framework. This framework automatically generates diverse types of open-ended QA pairs by leveraging domain-specific knowledge graphs 27 and evidence retrieved from PubMed (specialized) literature. Bio-KGR can be applied even in small-scale graph environments and ensures high quality dataset generation through literature-based filtering and reasoning difficulty adjustment. The generated QA pairs were evaluated and analyzed according to four criteria: question naturalness, answer appropriateness, reasoning difficulty, and dataset diversity through various LLM and expert evaluations. Experimental results showed that Bio-KGR demonstrated superior performance in most metrics not only in zero/few-shot settings but also when compared to existing manually constructed benchmarks. Furthermore, while evaluation results across models appeared consistently, the dataset was composed of problems with varying difficulty levels, demonstrating that Bio-KGR achieved high reliability and quality despite being completely automatically generated. These results show that biomedical QA benchmarks capable of systematically evaluating LLMs' high-dimensional reasoning abilities (in the biomedical domain) can be automatically constructed without manual work.
    번역하기

    Recently, with the advancement of large language models (LLMs), biomedical research has been accelerating, leading to a rapidly increasing demand for high-quality question-answering (QA) datasets. However, existing biomedical QA benchmarks rely on man...

    Recently, with the advancement of large language models (LLMs), biomedical research has been accelerating, leading to a rapidly increasing demand for high-quality question-answering (QA) datasets. However, existing biomedical QA benchmarks rely on manually created questions and fixed large-scale knowledge graphs, showing limitations in terms of currency, scalability, and efficiency. Since the biomedical field requires hierarchical reasoning capabilities and high-level domain expertise, with evidence based on specialized literature being essential, it is difficult to accurately evaluate LLMs' reasoning abilities using only superficial knowledge questions. To address these issues, we propose the Bio-KGR framework. This framework automatically generates diverse types of open-ended QA pairs by leveraging domain-specific knowledge graphs 27 and evidence retrieved from PubMed (specialized) literature. Bio-KGR can be applied even in small-scale graph environments and ensures high quality dataset generation through literature-based filtering and reasoning difficulty adjustment. The generated QA pairs were evaluated and analyzed according to four criteria: question naturalness, answer appropriateness, reasoning difficulty, and dataset diversity through various LLM and expert evaluations. Experimental results showed that Bio-KGR demonstrated superior performance in most metrics not only in zero/few-shot settings but also when compared to existing manually constructed benchmarks. Furthermore, while evaluation results across models appeared consistently, the dataset was composed of problems with varying difficulty levels, demonstrating that Bio-KGR achieved high reliability and quality despite being completely automatically generated. These results show that biomedical QA benchmarks capable of systematically evaluating LLMs' high-dimensional reasoning abilities (in the biomedical domain) can be automatically constructed without manual work.

    더보기

    목차 (Table of Contents)

    • 제1장 서론 1
    • 제2장 관련연구. 4
    • 제3장 방법론 6
    • 제1절 Bio-KGR 프레임워크 개요 6
    • 제2절 지식 그래프 기반 QA 구조 추출. 7
    • 제1장 서론 1
    • 제2장 관련연구. 4
    • 제3장 방법론 6
    • 제1절 Bio-KGR 프레임워크 개요 6
    • 제2절 지식 그래프 기반 QA 구조 추출. 7
    • 제3절 문헌 기반 질문-정답 생성 8
    • 제4절 LLM 기반 품질 평가 및 필터링. 9
    • 제4장 실험 설정 11
    • 제1절 데이터셋 품질 검증 11
    • 제2절 평가 모델. 11
    • 제3절 평가 지표. 11
    • 제5장 실험 결과 및 분석 13
    • 제1절 기존 지식 그래프 기반 QA 벤치마크와의 비교. 13
    • 제2절 프레임워크 구성 요소별 효과 분석 15
    • 제3절 평가 신뢰성 검증: Agreement Ratio 분석 16
    • 제4절 필터링 단계의 품질 보장 메커니즘 검증 17
    • 제5절 전문가 검증 18
    • 제6절 인간 평가자에 의한 난이도 평가 19
    • 제7절 데이터셋 규모의 영향 분석 19
    • 제8절 지식 그래프 간 적용 가능성 검증 21
    • 제6장 결론 22
    • 참고문헌 23
    • ABSTRACT 27
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼