본 연구는 거대 언어 모델(LLM)과 프롬프트 체이닝(Prompt Chaining) 기술을 활용하여 중학교 2학년 기하 단원의 수학 문항을 자동으로 생성하는 시스템을 구현하고, 생성된 문항의 교육적·심리측...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 연구는 거대 언어 모델(LLM)과 프롬프트 체이닝(Prompt Chaining) 기술을 활용하여 중학교 2학년 기하 단원의 수학 문항을 자동으로 생성하는 시스템을 구현하고, 생성된 문항의 교육적·심리측...
본 연구는 거대 언어 모델(LLM)과 프롬프트 체이닝(Prompt Chaining) 기술을 활용하여 중학교 2학년 기하 단원의 수학 문항을 자동으로 생성하는 시스템을 구현하고, 생성된 문항의 교육적·심리측정학적 타당성을 검증하는 데 목적이 있다. 기하 영역은 텍스트 발문과 시각적 도식 간의 엄밀한 논리적 정합성이 요구되는 분야로, 기존의 이미지 생성 모델들은 기하학적 제약 조건을 시각적으로 정확히 구현하는 데 한계를 보여 왔다.
이를 극복하기 위해 본 연구에서는 구글의 Gemini Pro 2.5 모델을 기반으로 ‘문항 본문 생성 – 도식 코드(TikZ) 생성 - 전문가 개입 순환(Human-in-the-Loop, HITL)’으로 이어지는 3단계 프롬프트 체이닝 절차를 설계하였다. 본 시스템은 기하 도식을 픽셀 이미지가 아닌 TikZ 코드로 생성하여 구조적 정확성을 확보하고, 생성 과정에서 발생하는 오류를 전문가가 식별하여 맞춤형 프롬프트로 수정하는 순환 구조를 핵심으로 한다.
연구 결과는 다음과 같다. 첫째, 문항 생성 과정에서 발생한 오류를 분석한 결과, 도형의 길이 비례나 위치 관계가 왜곡되는 기하학적 오류는 전체의 6.7%에 불과하여 코드 기반 생성 방식의 기술적 효용성이 입증되었다. 반면, 각도의 표시 방향이 반전되는 등의 맥락적 오류는 53.3%로 빈번하게 발생하였으나, 이는 HITL 과정을 통한 전문가의 개입으로 효과적으로 보정됨을 확인하였다.
둘째, 수학교육 전문가 3인을 대상으로 한 타당화 검사 결과, 최종 생성된 문항들은 교육과정 성취기준과의 부합성 및 구성 요소 간 일관성 측면에서 높은 내용 타당도(평균 CVR .689)를 확보한 것으로 평가되었다.
셋째, 문항반응이론(IRT)의 2모수 로지스틱 모형(2PLM)을 적용하여 심리측정학적 특성을 분석한 결과, 생성 문항의 변별도는 교과서 원본 문항과 거의 차이가 없었다. 다만, 난이도 측면에서는 생성 문항이 원본 문항보다 다소 낮게 형성되는 경향이 확인되었다.
본 연구는 프롬프트 체이닝과 인간-AI 협업 모델을 통해 기하 문항 자동생성의 기술적 난제를 해결할 수 있음을 실증하고, 생성된 문항이 실제 교육 현장의 평가 도구로서 기능할 수 있는 양호한 타당도를 지님을 확인하였다는 데 의의가 있다.
다국어 초록 (Multilingual Abstract)
The purpose of this study is to implement a system that automatically generates mathematics items for the geometry unit of the second grade of middle school using Large Language Models (LLM) and Prompt Chaining technology, and to verify the educationa...
The purpose of this study is to implement a system that automatically generates mathematics items for the geometry unit of the second grade of middle school using Large Language Models (LLM) and Prompt Chaining technology, and to verify the educational and psychometric validity of the generated items. The field of geometry requires rigorous logical consistency between text stems and visual diagrams; however, existing image generation models have shown limitations in visually implementing geometric constraints accurately.
To overcome this, this study designed a three-stage Prompt Chaining procedure based on Google’s Gemini Pro 2.5 model, consisting of "Item Stem Generation – Diagram Code (TikZ) Generation – Human-in-the-Loop (HITL)." The core of this system is to ensure structural accuracy by generating geometric diagrams as TikZ code rather than pixel-based images, and to employ a cyclic structure where experts identify errors occurring during the generation process and correct them through customized prompts.
The results of the study are as follows: First, an analysis of errors occurring during the item generation process revealed that geometric errors, such as distortions in length proportions or positional relationships, accounted for only 6.7% of the total, proving the technical utility of the code-based generation method. On the other hand, contextual errors, such as inverted angle markings, occurred frequently at 53.3%, but it was confirmed that these were effectively corrected through expert intervention via the HITL process.
Second, a validation test conducted by three mathematics education experts showed that the final generated items secured high content validity (Average CVR .689) in terms of alignment with curriculum achievement standards and consistency among item components.
Third, an analysis of psychometric characteristics applying the 2-Parameter Logistic Model (2PLM) of Item Response Theory (IRT) revealed that the discrimination of the generated items showed almost no difference from the original textbook items. However, in terms of difficulty, a tendency was confirmed where the generated items were formed at a slightly lower level than the original items.
This study is significant in demonstrating that technical challenges in the automatic generation of geometry items can be resolved through prompt chaining and human-AI collaboration models, and in confirming that the generated items possess sufficient validity to function as assessment tools in actual educational settings.
목차 (Table of Contents)