RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    소아수술실-회복실/중환자실 인계 요약지 생성 AI 모델 개발 = Development of an AI model for Generating Pediatric OR-to-Recovery Room/ICU Handover Documents

    한글로보기

    https://www.riss.kr/link?id=T17451604

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구는 소아 마취 데이터를 기반으로 회복실 및 중환자실 인계 요약지를 생성하는 거대 언어 모델(LLM)을 개발, LLM에 AI (Artificial Intelligence) 피드백 기반 강화학습(RLAIF)과 합성 데이터로 지도학습(SynSFT)을 적용하였을 때의 성능 변화를 비교하였다.
    2024년 1월 1일 ~ 2024년 12월 31일 사이 수술을 받은 소아 환자로 만 18세 이하의 환자 중 마취전상태평가 또는 마취기록 데이터가 있는 환자를 연구 대상으로 선정하여 마취전상태평가와 마취기록 데이터를 사용하였다. Meta-Llama-3.1-8B-Instruct와 Qwen3-8B 모델에 DPO, SimPO loss 함수를 적용하여 RLAIF를 진행하였다. 또한 인계 요약지 합성 데이터를 생성하여 이를 지도학습하는 SynSFT도 진행하였다. 평가는 정답 텍스트 없이 수행 가능한 LLM-as-Judge와 SCALE 지표를 사용하였으며, LLM-as-Judge에서는 Brevity & Relevance, Critical focus와 같이 2가지 기준을 잡고 평가를 진행하였다. 또한 LLM-as-Judge와 동일한 기준으로 human evaluation도 함께 진행하였다.
    그 결과, 모델마다의 차이가 존재하지만, SynSFT 또는 RLAIF를 적용하였을 때, 따로 학습을 진행하지 않은 원본 모델보다 인계 요약지 생성에 있어 개선된 성능을 보였다. Llama-3.1-8B-Instruct 모델의 경우, Base 모델의 LLM-as-Judge 두 항목의 총합 점수는 8.93임에 반해 SynSFT와 RLAIF를 둘 다 적용한 모델의 성능이 9.93으로 가장 높았다. 또한 Qwen3-8B는 Base 모델의 성능은 5.56이었으며, SynSFT를 적용한 모델의 성능이 9.74로 가장 높았다. Llama-3.1-8B-Instruct 모델에 대해서 human evaluation을 진행한 결과, 가장 점수를 높게 받은 모델은 SynSFT와 RLAIF를 둘 다 적용한 모델이었다. 각 모델들의 LLM-as-Judge의 점수 순위와 human evaluation 점수의 순위는 서로 유사한 경향성을 보였으며, 평가자 간의 일치도도 보통 수준 이상의 일치를 보였다.
    따라서, SynSFT와 RLAIF는 인계 요약지 생성 품질 향상 가능성을 지니며, 향후 의료진 피드백을 반영한 후속 연구를 통해 환자 인계 효율성에 기여할 것으로 기대된다.
    번역하기

    본 연구는 소아 마취 데이터를 기반으로 회복실 및 중환자실 인계 요약지를 생성하는 거대 언어 모델(LLM)을 개발, LLM에 AI (Artificial Intelligence) 피드백 기반 강화학습(RLAIF)과 합성 데이터로 지...

    본 연구는 소아 마취 데이터를 기반으로 회복실 및 중환자실 인계 요약지를 생성하는 거대 언어 모델(LLM)을 개발, LLM에 AI (Artificial Intelligence) 피드백 기반 강화학습(RLAIF)과 합성 데이터로 지도학습(SynSFT)을 적용하였을 때의 성능 변화를 비교하였다.
    2024년 1월 1일 ~ 2024년 12월 31일 사이 수술을 받은 소아 환자로 만 18세 이하의 환자 중 마취전상태평가 또는 마취기록 데이터가 있는 환자를 연구 대상으로 선정하여 마취전상태평가와 마취기록 데이터를 사용하였다. Meta-Llama-3.1-8B-Instruct와 Qwen3-8B 모델에 DPO, SimPO loss 함수를 적용하여 RLAIF를 진행하였다. 또한 인계 요약지 합성 데이터를 생성하여 이를 지도학습하는 SynSFT도 진행하였다. 평가는 정답 텍스트 없이 수행 가능한 LLM-as-Judge와 SCALE 지표를 사용하였으며, LLM-as-Judge에서는 Brevity & Relevance, Critical focus와 같이 2가지 기준을 잡고 평가를 진행하였다. 또한 LLM-as-Judge와 동일한 기준으로 human evaluation도 함께 진행하였다.
    그 결과, 모델마다의 차이가 존재하지만, SynSFT 또는 RLAIF를 적용하였을 때, 따로 학습을 진행하지 않은 원본 모델보다 인계 요약지 생성에 있어 개선된 성능을 보였다. Llama-3.1-8B-Instruct 모델의 경우, Base 모델의 LLM-as-Judge 두 항목의 총합 점수는 8.93임에 반해 SynSFT와 RLAIF를 둘 다 적용한 모델의 성능이 9.93으로 가장 높았다. 또한 Qwen3-8B는 Base 모델의 성능은 5.56이었으며, SynSFT를 적용한 모델의 성능이 9.74로 가장 높았다. Llama-3.1-8B-Instruct 모델에 대해서 human evaluation을 진행한 결과, 가장 점수를 높게 받은 모델은 SynSFT와 RLAIF를 둘 다 적용한 모델이었다. 각 모델들의 LLM-as-Judge의 점수 순위와 human evaluation 점수의 순위는 서로 유사한 경향성을 보였으며, 평가자 간의 일치도도 보통 수준 이상의 일치를 보였다.
    따라서, SynSFT와 RLAIF는 인계 요약지 생성 품질 향상 가능성을 지니며, 향후 의료진 피드백을 반영한 후속 연구를 통해 환자 인계 효율성에 기여할 것으로 기대된다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study developed a large language model (LLM) for generating postoperative handover documents for pediatric patients in recovery room and intensive care unit (ICU) and compared the performance changes when applying Reinforcement Learning with AI Feedback (RLAIF) and synthetic data supervised fine-tuning (SynSFT).
    Pediatric patients under 18 years of age who underwent surgery between January 1 and December 31, 2024, and had either pre-anesthetic assessment or anesthesia record data available were included in the study. These data were used as inputs, and RLAIF was applied to the Meta-Llama-3.1-8B-Instruct and Qwen3-8B models using DPO and SimPO loss functions for RLAIF, and additionally performed SynSFT using LLM-generated handover documents. Model performance was evaluated using LLM-as-Judge and SCALE metrics, both of which enable assessment without reference texts. In LLM-as-Judge evaluation, two criteria-Brevity & Relevance and Critical focus-were used. In addition, human evaluation was also conducted based on the same criteria as LLM-as-Judge.
    Overall, although the magnitude of the effect varied across models, applying SynSFT or RLAIF consistently improved handover document generation compared with the untrained base models. For Llama-3.1-8B-Instruct, the base model achieved a total LLM-as-Judge score of 8.93 across the two criteria, whereas the model trained with both SynSFT and RLAIF achieved the highest score (9.93). For Qwen3-8B, the base model scored 5.56, and the SynSFT model achieved the best performance (9.74). In the human evaluation of Llama-3.1-8B-Instruct, the highest-scoring model was again the one trained with both SynSFT and RLAIF. The ranking of models by LLM-as-Judge scores showed a similar trend to the ranking from human evaluation, and inter-rater agreement was at least moderate.
    These results suggest that SynSFT and RLAIF has potential to enhance the quality of handover document generation. Future work will incorporate clinicians’ feedback to further improve LLM performance and contribute to more efficient clinical handover documents.
    번역하기

    This study developed a large language model (LLM) for generating postoperative handover documents for pediatric patients in recovery room and intensive care unit (ICU) and compared the performance changes when applying Reinforcement Learning with AI F...

    This study developed a large language model (LLM) for generating postoperative handover documents for pediatric patients in recovery room and intensive care unit (ICU) and compared the performance changes when applying Reinforcement Learning with AI Feedback (RLAIF) and synthetic data supervised fine-tuning (SynSFT).
    Pediatric patients under 18 years of age who underwent surgery between January 1 and December 31, 2024, and had either pre-anesthetic assessment or anesthesia record data available were included in the study. These data were used as inputs, and RLAIF was applied to the Meta-Llama-3.1-8B-Instruct and Qwen3-8B models using DPO and SimPO loss functions for RLAIF, and additionally performed SynSFT using LLM-generated handover documents. Model performance was evaluated using LLM-as-Judge and SCALE metrics, both of which enable assessment without reference texts. In LLM-as-Judge evaluation, two criteria-Brevity & Relevance and Critical focus-were used. In addition, human evaluation was also conducted based on the same criteria as LLM-as-Judge.
    Overall, although the magnitude of the effect varied across models, applying SynSFT or RLAIF consistently improved handover document generation compared with the untrained base models. For Llama-3.1-8B-Instruct, the base model achieved a total LLM-as-Judge score of 8.93 across the two criteria, whereas the model trained with both SynSFT and RLAIF achieved the highest score (9.93). For Qwen3-8B, the base model scored 5.56, and the SynSFT model achieved the best performance (9.74). In the human evaluation of Llama-3.1-8B-Instruct, the highest-scoring model was again the one trained with both SynSFT and RLAIF. The ranking of models by LLM-as-Judge scores showed a similar trend to the ranking from human evaluation, and inter-rater agreement was at least moderate.
    These results suggest that SynSFT and RLAIF has potential to enhance the quality of handover document generation. Future work will incorporate clinicians’ feedback to further improve LLM performance and contribute to more efficient clinical handover documents.

    더보기

    목차 (Table of Contents)

    • 제 1 장 서론 1
    • 제 1 절 연구의 배경 1
    • 제 2 절 연구 가설 및 목적 2
    • 제 2 장 이론적 배경 및 관련 연구 3
    • 제 1 장 서론 1
    • 제 1 절 연구의 배경 1
    • 제 2 절 연구 가설 및 목적 2
    • 제 2 장 이론적 배경 및 관련 연구 3
    • 제 1 절 인계 요약지 3
    • 제 2 절 거대 언어 모델 3
    • 제 3 절 강화학습과 선호도 기반 정렬 학습 4
    • 제 4 절 합성 데이터 기반 지도학습 5
    • 제 3 장 연구 방법 6
    • 제 1 절 데이터 추출 및 전처리 6
    • 제 2 절 모델 7
    • 제 3 절 평가 방법 17
    • 제 4 절 실험 환경 26
    • 제 4 장 연구 결과 27
    • 제 1 절 LLaMA 모델 성능 비교 27
    • 제 2 절 Qwen 모델 성능 비교 28
    • 제 3 절 Human Evaluation 29
    • 제 5 장 고찰 및 결론 34
    • 제 1 절 LLM 종류에 따른 성능 차이 34
    • 제 2 절 RLAIF 및 SynSFT의 필요성 34
    • 제 3 절 Human Evaluation과 LLM-as-Judge 35
    • 제 4 절 임상 환경 적용 시나리오 36
    • 제 5 절 한계점 38
    • 제 6 절 결론 38
    • 참고문헌 39
    • Abstract 45
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼