최근 대규모 언어 모델(Large Language Model; LLM)의 발전과 함께 바이오메디컬 분야 연구가 가속화되면서 고품질 질문-답변(QA) 데이터셋에 대한 수요가 빠르게 늘고 있다. 그러나 기존 벤치마크는...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
최근 대규모 언어 모델(Large Language Model; LLM)의 발전과 함께 바이오메디컬 분야 연구가 가속화되면서 고품질 질문-답변(QA) 데이터셋에 대한 수요가 빠르게 늘고 있다. 그러나 기존 벤치마크는...
최근 대규모 언어 모델(Large Language Model; LLM)의 발전과 함께 바이오메디컬 분야 연구가 가속화되면서 고품질 질문-답변(QA) 데이터셋에 대한 수요가 빠르게 늘고 있다. 그러나 기존 벤치마크는 수작업 질문과 고정된 지식 그래프에 의존해 최신성·확장성·효율성에 한계를 보인다. 특히, 이 분야는 계층적 추론 능력과 전문 지식을 요구하며 전문 문헌 기반의 근거가 필요하기 때문에 표면적 지식 질문 만으로 LLM의 추론 능력을 정확히 평가하기 어렵다. 이러한 한계를 해결하기 위해 Bio-KGR 프레임워크를 제안한다. 이 프레임워크는 도메인 특화 지식 그래프와 검색 증강 생성을 통해 PubMed 문헌에서 검색된 근거를 활용하여 다양한 유형의 개방형 QA 쌍을 자동으로 생성한다. 소규모 그래프에도 적용 가능하며, 문헌 기반 필터링과 난이도 조절을 통해 데이터 품질을 확보한다. 생성된 QA 쌍은 다양한 LLM 및 전문가 평가를 통하여 질문 자연스러움, 답변 적절성, 추론 난이도, 데이터셋 다양성이라는 네 가지 기준으로 분석되었다. 실험 결과, Bio-KGR는 제로/퓨샷 설정 및 기존 수작업 벤치마크와 비교했을 때도 대부분의 지표에서 우수한 성능을 보였다. 또한 모델 간 일관된 평가 결과와 다양한 난이도 구성으로 Bio-KGR는 자동 생성임에도 불구하고 높은 신뢰성과 품질을 입증했다. 이러한 결과는 LLM의 고차원적 추론 능력을 체계적으로 평가할 수 있는 바이오메디컬 QA 벤치마크를 수작업 없이 자동으로 구축할 수 있음을 보여준다.
다국어 초록 (Multilingual Abstract)
Recently, with the advancement of large language models (LLMs), biomedical research has been accelerating, leading to a rapidly increasing demand for high-quality question-answering (QA) datasets. However, existing biomedical QA benchmarks rely on man...
Recently, with the advancement of large language models (LLMs), biomedical research has been accelerating, leading to a rapidly increasing demand for high-quality question-answering (QA) datasets. However, existing biomedical QA benchmarks rely on manually created questions and fixed large-scale knowledge graphs, showing limitations in terms of currency, scalability, and efficiency. Since the biomedical field requires hierarchical reasoning capabilities and high-level domain expertise, with evidence based on specialized literature being essential, it is difficult to accurately evaluate LLMs' reasoning abilities using only superficial knowledge questions. To address these issues, we propose the Bio-KGR framework. This framework automatically generates diverse types of open-ended QA pairs by leveraging domain-specific knowledge graphs 27 and evidence retrieved from PubMed (specialized) literature. Bio-KGR can be applied even in small-scale graph environments and ensures high quality dataset generation through literature-based filtering and reasoning difficulty adjustment. The generated QA pairs were evaluated and analyzed according to four criteria: question naturalness, answer appropriateness, reasoning difficulty, and dataset diversity through various LLM and expert evaluations. Experimental results showed that Bio-KGR demonstrated superior performance in most metrics not only in zero/few-shot settings but also when compared to existing manually constructed benchmarks. Furthermore, while evaluation results across models appeared consistently, the dataset was composed of problems with varying difficulty levels, demonstrating that Bio-KGR achieved high reliability and quality despite being completely automatically generated. These results show that biomedical QA benchmarks capable of systematically evaluating LLMs' high-dimensional reasoning abilities (in the biomedical domain) can be automatically constructed without manual work.
목차 (Table of Contents)