RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    프롬프트 체이닝과 인간 개입 순환 구조를 활용한 거대 언어 모델 기반 문항 생성 방법 연구 : 중학교 2학년 기하 단원을 중심으로 = A Study on Large Language Model-Based Item Generation Methods Using Prompt Chaining and Human-in-the-Loop Structures: Focusing on the Middle School Geometry

    한글로보기

    https://www.riss.kr/link?id=T17451950

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구는 거대 언어 모델(LLM)과 프롬프트 체이닝(Prompt Chaining) 기술을 활용하여 중학교 2학년 기하 단원의 수학 문항을 자동으로 생성하는 시스템을 구현하고, 생성된 문항의 교육적·심리측정학적 타당성을 검증하는 데 목적이 있다. 기하 영역은 텍스트 발문과 시각적 도식 간의 엄밀한 논리적 정합성이 요구되는 분야로, 기존의 이미지 생성 모델들은 기하학적 제약 조건을 시각적으로 정확히 구현하는 데 한계를 보여 왔다.
    이를 극복하기 위해 본 연구에서는 구글의 Gemini Pro 2.5 모델을 기반으로 ‘문항 본문 생성 – 도식 코드(TikZ) 생성 - 전문가 개입 순환(Human-in-the-Loop, HITL)’으로 이어지는 3단계 프롬프트 체이닝 절차를 설계하였다. 본 시스템은 기하 도식을 픽셀 이미지가 아닌 TikZ 코드로 생성하여 구조적 정확성을 확보하고, 생성 과정에서 발생하는 오류를 전문가가 식별하여 맞춤형 프롬프트로 수정하는 순환 구조를 핵심으로 한다.
    연구 결과는 다음과 같다. 첫째, 문항 생성 과정에서 발생한 오류를 분석한 결과, 도형의 길이 비례나 위치 관계가 왜곡되는 기하학적 오류는 전체의 6.7%에 불과하여 코드 기반 생성 방식의 기술적 효용성이 입증되었다. 반면, 각도의 표시 방향이 반전되는 등의 맥락적 오류는 53.3%로 빈번하게 발생하였으나, 이는 HITL 과정을 통한 전문가의 개입으로 효과적으로 보정됨을 확인하였다.
    둘째, 수학교육 전문가 3인을 대상으로 한 타당화 검사 결과, 최종 생성된 문항들은 교육과정 성취기준과의 부합성 및 구성 요소 간 일관성 측면에서 높은 내용 타당도(평균 CVR .689)를 확보한 것으로 평가되었다.
    셋째, 문항반응이론(IRT)의 2모수 로지스틱 모형(2PLM)을 적용하여 심리측정학적 특성을 분석한 결과, 생성 문항의 변별도는 교과서 원본 문항과 거의 차이가 없었다. 다만, 난이도 측면에서는 생성 문항이 원본 문항보다 다소 낮게 형성되는 경향이 확인되었다.
    본 연구는 프롬프트 체이닝과 인간-AI 협업 모델을 통해 기하 문항 자동생성의 기술적 난제를 해결할 수 있음을 실증하고, 생성된 문항이 실제 교육 현장의 평가 도구로서 기능할 수 있는 양호한 타당도를 지님을 확인하였다는 데 의의가 있다.
    번역하기

    본 연구는 거대 언어 모델(LLM)과 프롬프트 체이닝(Prompt Chaining) 기술을 활용하여 중학교 2학년 기하 단원의 수학 문항을 자동으로 생성하는 시스템을 구현하고, 생성된 문항의 교육적·심리측...

    본 연구는 거대 언어 모델(LLM)과 프롬프트 체이닝(Prompt Chaining) 기술을 활용하여 중학교 2학년 기하 단원의 수학 문항을 자동으로 생성하는 시스템을 구현하고, 생성된 문항의 교육적·심리측정학적 타당성을 검증하는 데 목적이 있다. 기하 영역은 텍스트 발문과 시각적 도식 간의 엄밀한 논리적 정합성이 요구되는 분야로, 기존의 이미지 생성 모델들은 기하학적 제약 조건을 시각적으로 정확히 구현하는 데 한계를 보여 왔다.
    이를 극복하기 위해 본 연구에서는 구글의 Gemini Pro 2.5 모델을 기반으로 ‘문항 본문 생성 – 도식 코드(TikZ) 생성 - 전문가 개입 순환(Human-in-the-Loop, HITL)’으로 이어지는 3단계 프롬프트 체이닝 절차를 설계하였다. 본 시스템은 기하 도식을 픽셀 이미지가 아닌 TikZ 코드로 생성하여 구조적 정확성을 확보하고, 생성 과정에서 발생하는 오류를 전문가가 식별하여 맞춤형 프롬프트로 수정하는 순환 구조를 핵심으로 한다.
    연구 결과는 다음과 같다. 첫째, 문항 생성 과정에서 발생한 오류를 분석한 결과, 도형의 길이 비례나 위치 관계가 왜곡되는 기하학적 오류는 전체의 6.7%에 불과하여 코드 기반 생성 방식의 기술적 효용성이 입증되었다. 반면, 각도의 표시 방향이 반전되는 등의 맥락적 오류는 53.3%로 빈번하게 발생하였으나, 이는 HITL 과정을 통한 전문가의 개입으로 효과적으로 보정됨을 확인하였다.
    둘째, 수학교육 전문가 3인을 대상으로 한 타당화 검사 결과, 최종 생성된 문항들은 교육과정 성취기준과의 부합성 및 구성 요소 간 일관성 측면에서 높은 내용 타당도(평균 CVR .689)를 확보한 것으로 평가되었다.
    셋째, 문항반응이론(IRT)의 2모수 로지스틱 모형(2PLM)을 적용하여 심리측정학적 특성을 분석한 결과, 생성 문항의 변별도는 교과서 원본 문항과 거의 차이가 없었다. 다만, 난이도 측면에서는 생성 문항이 원본 문항보다 다소 낮게 형성되는 경향이 확인되었다.
    본 연구는 프롬프트 체이닝과 인간-AI 협업 모델을 통해 기하 문항 자동생성의 기술적 난제를 해결할 수 있음을 실증하고, 생성된 문항이 실제 교육 현장의 평가 도구로서 기능할 수 있는 양호한 타당도를 지님을 확인하였다는 데 의의가 있다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The purpose of this study is to implement a system that automatically generates mathematics items for the geometry unit of the second grade of middle school using Large Language Models (LLM) and Prompt Chaining technology, and to verify the educational and psychometric validity of the generated items. The field of geometry requires rigorous logical consistency between text stems and visual diagrams; however, existing image generation models have shown limitations in visually implementing geometric constraints accurately.
    To overcome this, this study designed a three-stage Prompt Chaining procedure based on Google’s Gemini Pro 2.5 model, consisting of "Item Stem Generation – Diagram Code (TikZ) Generation – Human-in-the-Loop (HITL)." The core of this system is to ensure structural accuracy by generating geometric diagrams as TikZ code rather than pixel-based images, and to employ a cyclic structure where experts identify errors occurring during the generation process and correct them through customized prompts.
    The results of the study are as follows: First, an analysis of errors occurring during the item generation process revealed that geometric errors, such as distortions in length proportions or positional relationships, accounted for only 6.7% of the total, proving the technical utility of the code-based generation method. On the other hand, contextual errors, such as inverted angle markings, occurred frequently at 53.3%, but it was confirmed that these were effectively corrected through expert intervention via the HITL process.
    Second, a validation test conducted by three mathematics education experts showed that the final generated items secured high content validity (Average CVR .689) in terms of alignment with curriculum achievement standards and consistency among item components.
    Third, an analysis of psychometric characteristics applying the 2-Parameter Logistic Model (2PLM) of Item Response Theory (IRT) revealed that the discrimination of the generated items showed almost no difference from the original textbook items. However, in terms of difficulty, a tendency was confirmed where the generated items were formed at a slightly lower level than the original items.
    This study is significant in demonstrating that technical challenges in the automatic generation of geometry items can be resolved through prompt chaining and human-AI collaboration models, and in confirming that the generated items possess sufficient validity to function as assessment tools in actual educational settings.
    번역하기

    The purpose of this study is to implement a system that automatically generates mathematics items for the geometry unit of the second grade of middle school using Large Language Models (LLM) and Prompt Chaining technology, and to verify the educationa...

    The purpose of this study is to implement a system that automatically generates mathematics items for the geometry unit of the second grade of middle school using Large Language Models (LLM) and Prompt Chaining technology, and to verify the educational and psychometric validity of the generated items. The field of geometry requires rigorous logical consistency between text stems and visual diagrams; however, existing image generation models have shown limitations in visually implementing geometric constraints accurately.
    To overcome this, this study designed a three-stage Prompt Chaining procedure based on Google’s Gemini Pro 2.5 model, consisting of "Item Stem Generation – Diagram Code (TikZ) Generation – Human-in-the-Loop (HITL)." The core of this system is to ensure structural accuracy by generating geometric diagrams as TikZ code rather than pixel-based images, and to employ a cyclic structure where experts identify errors occurring during the generation process and correct them through customized prompts.
    The results of the study are as follows: First, an analysis of errors occurring during the item generation process revealed that geometric errors, such as distortions in length proportions or positional relationships, accounted for only 6.7% of the total, proving the technical utility of the code-based generation method. On the other hand, contextual errors, such as inverted angle markings, occurred frequently at 53.3%, but it was confirmed that these were effectively corrected through expert intervention via the HITL process.
    Second, a validation test conducted by three mathematics education experts showed that the final generated items secured high content validity (Average CVR .689) in terms of alignment with curriculum achievement standards and consistency among item components.
    Third, an analysis of psychometric characteristics applying the 2-Parameter Logistic Model (2PLM) of Item Response Theory (IRT) revealed that the discrimination of the generated items showed almost no difference from the original textbook items. However, in terms of difficulty, a tendency was confirmed where the generated items were formed at a slightly lower level than the original items.
    This study is significant in demonstrating that technical challenges in the automatic generation of geometry items can be resolved through prompt chaining and human-AI collaboration models, and in confirming that the generated items possess sufficient validity to function as assessment tools in actual educational settings.

    더보기

    목차 (Table of Contents)

    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구 질문 5
    • Ⅱ. 이론적 배경 6
    • 1. 자동문항생성 6
    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구 질문 5
    • Ⅱ. 이론적 배경 6
    • 1. 자동문항생성 6
    • 2. 프롬프트 엔지니어링 11
    • 3. 문항반응이론 13
    • Ⅲ. 연구 방법 19
    • 1. 프롬프트 체이닝 및 인간 개입 순환 기반 문항 생성 19
    • 2. 전문가 검토 및 내용 타당도 검증 28
    • 3. 문항 반응 데이터 수집 방법 30
    • 4. 문항반응이론 기반 문항 분석 33
    • Ⅳ. 연구 결과 35
    • 1. 프롬프트 체이닝 기반 문항 생성 과정 및 오류 분석 35
    • 2. 전문가 검토를 통한 문항 타당성 검증 39
    • 3. 문항반응이론 기반 정량적 분석 결과 41
    • Ⅴ. 결론 및 제언 46
    • 1. 결론 46
    • 2. 논의 47
    • 참고문헌 49
    • 부록 55
    • Abstract 102
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼