본 연구는 중·고등학교 수학교육에서 개인 맞춤형 AI 챗봇 제작 및 수학 문제해결 수업을 위한 설계원리 및 상세지침을 개발하고, 타당화하는 것을 목적으로 한다. 생성형 인공지능 기술이 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 연구는 중·고등학교 수학교육에서 개인 맞춤형 AI 챗봇 제작 및 수학 문제해결 수업을 위한 설계원리 및 상세지침을 개발하고, 타당화하는 것을 목적으로 한다. 생성형 인공지능 기술이 ...
본 연구는 중·고등학교 수학교육에서 개인 맞춤형 AI 챗봇 제작 및 수학 문제해결 수업을 위한 설계원리 및 상세지침을 개발하고, 타당화하는 것을 목적으로 한다. 생성형 인공지능 기술이 빠르게 고도화되는 가운데, ChatGPT를 비롯하여 Claude, Gemini, Grok 등 대규모 언어모델 기반 챗봇이 교육 영역에도 적용되기 시작하였다. 국내에서는 2022 개정 교육과정 고시와 함께 교육부가 AI 디지털교과서 및 AI 보조교사 활용 방안을 제시하면서 인공지능 기반 맞춤형 교육에 대한 정책적 관심이 높아지고 있다. 그럼에도 중·고등학교 수학 교육 맥락에서 교사가 학생 개개인에게 맞춘 AI 챗봇을 설계하고 수업에 체계적으로 적용하기 위한 원리를 다룬 연구는 찾아보기 어렵다. 이러한 배경에서 본 연구는 Richey와 Klein(2007)의 설계·개발 연구 방법론 중 유형2-모형 연구에 따라 현장에서 일반화하여 활용할 수 있는 설계원리와 상세지침을 개발하고자 하였다. 연구는 크게 초기 설계원리 개발, 내적 타당화, 맞춤형 AI 챗봇 플랫폼 개발, 외적 타당화, 최종 설계원리 도출의 다섯 단계로 진행되었다.
먼저 선행문헌 검토, 경험적 탐색, 전문가 면담을 통해 초기 설계원리를 개발하였다. 선행문헌은 ‘수학적 문제해결, 개인 맞춤형 수업, 생성형 인공지능 챗봇’의 세 가지 영역을 중심으로 검토하였고, Polya(1945)와 Schoenfeld(1985)의 문제해결 이론, 개인 맞춤형 교육의 개념 및 구성요소, 프롬프트 엔지니어링 기법 등에 관한 이론적 기반을 탐색하였다. 경험적 탐색으로는 연구자가 직접 Google Apps Script를 활용하여 맞춤형 AI 챗봇 플랫폼의 프로토타입을 개발하고, 2025년 3월부터 약 2개월간 인천 소재 중학교 2학년 학생 124명을 대상으로 파일럿 테스트를 실시하였다. 이 과정에서 학생별 학습 수준, 태도, 관심사를 반영한 개별화 시스템 프롬프트를 구성하고, 교과서 5개 단원의 80개 문항에 대한 문제해결 프롬프트를 설계하여 실제 수업에 적용함으로써 현장 데이터를 수집하였다. 이러한 과정을 통해 13개의 설계원리와 27개의 상세지침으로 구성된 초기 설계원리를 도출하였고, 이를 수학교육 전문가 3인, 교육공학 전문가 3인 총 6명의 전문가 패널을 대상으로 2차에 걸친 타당화를 진행하였다. 나일주, 정현미(2001)가 제시한 타당성, 설명력, 유용성, 보편성, 이해도의 5가지 준거를 4점 리커트 척도로 타당화를 진행하였다. 1차 타당화 결과 전문가들은 설계원리의 구성 체계가 챗봇 설계와 수업 설계로 혼재되어 일관성이 부족하다는 점, 수학교육의 특수성이 원리 자체에 충분히 반영되지 않고 예시에만 나타난다는 점 등을 지적하였다. 이러한 피드백을 바탕으로 설계원리를 전면적으로 재구조화하여 초기의 5개 구성요소를 ‘맞춤형 AI 챗봇, 교수자, 학습자, 학습환경’의 4개 영역으로 재분류하여 설계 주체와 기능에 따른 명확한 체계를 구축하는 등 11개의 설계원리와 32개의 상세지침으로 2차 설계원리를 개발하였다. 2차 타당화 결과 타당성, 유용성, 보편성 영역에서는 전문가 전원이 4점(매우 그렇다)을 부여하여 평균 4.00점을 기록하였으며, 설명력(평균 3.83점)과 이해도(평균 3.50점)도 1차 대비 크게 향상되었다. 내용타당도 지수(CVI)는 모든 영역에서 1.0으로 나타나 설계원리의 타당성이 매우 높음을 확인하였으며, 채점자 간 일치도 지수(IRA) 역시 1.0으로 전문가 평가의 신뢰도가 높았다. 개별 설계원리별로는 11개 원리 모두 평균 3.67점 이상을 기록하였다. 또한, 교사들이 맞춤형 AI 챗봇을 쉽게 설계할 수 있는 플랫폼을 개발하였다. 챗봇 플랫폼은 Google Cloud Platform의 Cloud Run을 활용한 서버리스 아키텍처로 구축되었으며, 프론트엔드는 Svelte, 백엔드는 FastAPI를 사용하여 개발하였다. 교사는 실시간 대시보드를 통해 학생별 대화 내용과 진행 상황을 모니터링할 수 있다.
개발된 설계원리의 현장 적용 가능성을 검증하기 위해 12월에 외적 타당화를 실시하였다. 경기도 소재 중학교 1개교(3학년 30명)와 경상북도 소재 고등학교 1개교(2학년 30명)를 대상으로, 교사가 설계한 맞춤형 챗봇을 활용하여 3차시 수업을 진행하고 학습자와 교수자의 반응을 수집 및 분석하였다. 학습자들은 즉시적 접근성, 개별화된 설명, 정서적 편안함 등을 강점으로 인식하였으며, 수식 렌더링 오류, 학년 수준 초과 개념 사용, 비계 유연성 부족 등을 개선점으로 제시하였다. 교수자들은 설계원리의 타당성과 유용성을 긍정적으로 평가하면서, 하이퍼 파라미터 설명 보완, 변형문제를 활용한 연습, 시스템 프롬프트 개선 자동화 등을 제안하였다. 이러한 외적 타당화 결과를 반영하여 비계 유연화, 수식 렌더링 지침 강화, 변형 문제 제공 기능 등을 보완한 최종 설계원리를 도출하였다.
본 연구는 다음과 같은 제한점을 가지고 있다. 첫째, 외적 타당화를 중고등학교 교수자 2명, 학습자 60명을 대상으로 실시하여 연구 대상의 수와 범위가 제한적이다. 다양한 학교급, 지역, 학습자 특성을 가진 집단을 대상으로 한 현장 적용 연구가 필요하다. 둘째, 본 연구는 수학 문제해결 영역에 초점을 맞추었으나 개념 학습, 탐구 활동, 프로젝트 수업 등 다른 유형의 수학 수업이나 타 교과로의 확장을 위해 추가 연구가 필요하다. 셋째, 외적 타당화 과정에서 통제 집단을 설정한 준실험 설계나 사전-사후 검사를 통한 정량적 효과 검증은 실시하지 않았다. 학습자의 대화 데이터와 반응을 질적으로 분석하였으며, 설계원리 적용에 따른 학업 성취도 변화 등 양적 효과 검증은 후속 연구에서 다루어질 필요가 있다.
본 연구는 다음과 같은 의의를 갖는다. 첫째, 생성형 AI 챗봇을 활용한 개인 맞춤형 수학 문제해결 수업의 체계적인 설계원리를 제시함으로써 이론적 기반을 마련하였다. 기존 연구들이 범용 AI의 활용에 그쳤다면, 본 연구는 교사가 학생 맞춤형 AI 챗봇을 직접 설계하고, 설계한 챗봇과 학생이 상호작용하면서 AI 보조교사로서 활용할 수 있는 가능성을 구체적인 지침과 함께 제시하였다. 둘째, 챗봇 설계와 수업 설계를 통합한 포괄적 원리를 개발함으로써 AI 시대의 새로운 교수설계 방식을 제안하였다. 셋째, 파일럿 테스트를 통한 경험적 데이터, 전문가 타당화를 통한 이론적 정교화, 외적 타당화를 통한 현장 검증을 거쳐 현장 적용 가능성과 학술적 타당성을 동시에 확보하였다.
다국어 초록 (Multilingual Abstract)
The purpose of this study is to develop and validate design principles and detailed guidelines for creating personalized AI chatbots and implementing mathematical problem-solving instruction in secondary school mathematics education. Amid the rapid ad...
The purpose of this study is to develop and validate design principles and detailed guidelines for creating personalized AI chatbots and implementing mathematical problem-solving instruction in secondary school mathematics education. Amid the rapid advancement of generative artificial intelligence, large language model (LLM)-based chatbots such as ChatGPT, Claude, Gemini, and Grok have emerged, driving innovative changes in the field of education. In particular, the 2022 Revised National Curriculum and the Ministry of Education's digital-based education innovation policy emphasize the realization of personalized education utilizing artificial intelligence, including the introduction of AI Digital Textbooks (AIDT) and the use of AI teaching assistants. However, despite these policy demands, research on design principles that enable teachers to directly design personalized AI chatbots and systematically apply them in secondary school mathematics classrooms remains scarce. Therefore, this study applied Type 2 (model research) of the Design and Development Research methodology proposed by Richey and Klein (2007) to develop generalizable design principles and detailed guidelines. The research was conducted in five stages: initial design principle development, internal validation, personalized AI chatbot platform development, external validation, and final design principle derivation.
First, initial design principles were developed through literature review, empirical exploration, and expert interviews. The literature review focused on three areas: mathematical problem-solving, personalized instruction, and generative AI chatbots, exploring the theoretical foundations of Polya's (1945) and Schoenfeld's (1985) problem-solving theories, concepts and components of personalized education, and prompt engineering techniques. For empirical exploration, the researcher directly developed a prototype of a personalized AI chatbot platform using Google Apps Script and conducted a pilot test with 124 second-year middle school students in Incheon over approximately two months starting from March 2025. During this process, individualized system prompts reflecting each student's learning level, attitude, and interests were constructed, and problem-solving prompts for 80 questions across five textbook units were designed and applied in actual classes to collect field data. Through this process, initial design principles consisting of 13 principles and 27 detailed guidelines were derived, and two rounds of validation were conducted with an expert panel of six members, including three mathematics education experts and three educational technology experts. Validation was conducted using a 4-point Likert scale based on five criteria presented by Nah and Chung (2001): validity, explanatory power, usefulness, universality, and comprehensibility. In the first validation, experts pointed out that the structural system of design principles lacked consistency due to the mixture of chatbot design and instructional design, and that the specificity of mathematics education was not sufficiently reflected in the principles themselves but only appeared in examples. Based on this feedback, the design principles were completely restructured, reclassifying the initial five components into four areas—'Personalized AI Chatbot, Instructor, Learner, and Learning Environment'—to establish a clear system according to design agents and functions, resulting in 11 design principles and 32 detailed guidelines for the second version. In the second validation, all experts gave 4 points (strongly agree) in the areas of validity, usefulness, and universality, recording an average of 4.00 points, while explanatory power (average 3.83 points) and comprehensibility (average 3.50 points) also improved significantly compared to the first round. The Content Validity Index (CVI) was 1.0 in all areas, confirming very high validity of the design principles, and the Inter-Rater Agreement (IRA) was also 1.0, indicating high reliability of expert evaluation. All 11 individual design principles recorded an average of 3.67 points or higher.
Additionally, a platform was developed to enable teachers to easily design personalized AI chatbots. The platform was built with a serverless architecture utilizing Cloud Run on Google Cloud Platform (GCP), with Svelte for the frontend and FastAPI for the backend. Teachers can monitor individual student conversations and progress through a real-time dashboard.
External validation was conducted in December to verify the field applicability of the developed design principles. Three sessions of instruction were conducted using teacher-designed personalized chatbots at one middle school in Gyeonggi Province (30 third-year students) and one high school in Gyeongsangbuk Province (30 second-year students), and learner and instructor responses were collected and analyzed. Learners recognized immediate accessibility, individualized explanations, and emotional comfort as strengths, while suggesting improvements such as rendering errors in mathematical formulas, use of concepts beyond grade level, and lack of flexible scaffolding. Instructors positively evaluated the validity and usefulness of the design principles while suggesting supplementary explanations for hyperparameters, functions for generating variant problems, and automation of system prompt improvement. Reflecting these external validation results, the final design principles were derived with enhancements including flexible scaffolding, grade-level concept control, strengthened formula rendering guidelines, and variant problem provision functions.
This study has the following limitations. First, although external validation was conducted at one middle school and one high school, the number and scope of research subjects were limited, presenting limitations for generalizing the design principles. Additional field application studies with diverse school levels, regions, and learner characteristics are needed. Second, this study focused on the mathematical problem-solving domain, and additional research is needed for the possibility of extension to other types of mathematics instruction such as concept learning, inquiry activities, and project-based learning, as well as to other subjects. Third, the external validation process did not employ quasi-experimental designs with control groups or quantitative effect verification through pre-post tests. Learner conversation data and responses were analyzed qualitatively, and quantitative effect verification such as changes in academic achievement resulting from the application of design principles needs to be addressed in follow-up studies.
This study has the following significance. First, it established a theoretical foundation by presenting systematic design principles for personalized mathematical problem-solving instruction using generative AI chatbots. While existing studies remained at the level of utilizing general-purpose AI, this study presented the possibility of teachers directly designing student-personalized AI chatbots and utilizing them as AI teaching assistants through student interaction, along with specific guidelines. Second, by developing comprehensive principles integrating chatbot design and instructional design, it proposed a new instructional design approach for the AI era. Third, through empirical data from pilot testing, theoretical refinement through expert validation, and field verification through external validation, both field applicability and academic validity were simultaneously secured.The purpose of this study is to develop and validate design principles and detailed guidelines for creating personalized AI chatbots and implementing mathematical problem-solving instruction in secondary school mathematics education. Amid the rapid advancement of generative artificial intelligence, large language model (LLM)-based chatbots such as ChatGPT, Claude, Gemini, and Grok have emerged, driving innovative changes in the field of education. In particular, the 2022 Revised National Curriculum and the Ministry of Education's digital-based education innovation policy emphasize the realization of personalized education utilizing artificial intelligence, including the introduction of AI Digital Textbooks (AIDT) and the use of AI teaching assistants. However, despite these policy demands, research on design principles that enable teachers to directly design personalized AI chatbots and systematically apply them in secondary school mathematics classrooms remains scarce. Therefore, this study applied Type 2 (model research) of the Design and Development Research methodology proposed by Richey and Klein (2007) to develop generalizable design principles and detailed guidelines. The research was conducted in five stages: initial design principle development, internal validation, personalized AI chatbot platform development, external validation, and final design principle derivation.
First, initial design principles were developed through literature review, empirical exploration, and expert interviews. The literature review focused on three areas: mathematical problem-solving, personalized instruction, and generative AI chatbots, exploring the theoretical foundations of Polya's (1945) and Schoenfeld's (1985) problem-solving theories, concepts and components of personalized education, and prompt engineering techniques. For empirical exploration, the researcher directly developed a prototype of a personalized AI chatbot platform using Google Apps Script and conducted a pilot test with 124 second-year middle school students in Incheon over approximately two months starting from March 2025. During this process, individualized system prompts reflecting each student's learning level, attitude, and interests were constructed, and problem-solving prompts for 80 questions across five textbook units were designed and applied in actual classes to collect field data. Through this process, initial design principles consisting of 13 principles and 27 detailed guidelines were derived, and two rounds of validation were conducted with an expert panel of six members, including three mathematics education experts and three educational technology experts. Validation was conducted using a 4-point Likert scale based on five criteria presented by Nah and Chung (2001): validity, explanatory power, usefulness, universality, and comprehensibility. In the first validation, experts pointed out that the structural system of design principles lacked consistency due to the mixture of chatbot design and instructional design, and that the specificity of mathematics education was not sufficiently reflected in the principles themselves but only appeared in examples. Based on this feedback, the design principles were completely restructured, reclassifying the initial five components into four areas—'Personalized AI Chatbot, Instructor, Learner, and Learning Environment'—to establish a clear system according to design agents and functions, resulting in 11 design principles and 32 detailed guidelines for the second version. In the second validation, all experts gave 4 points (strongly agree) in the areas of validity, usefulness, and universality, recording an average of 4.00 points, while explanatory power (average 3.83 points) and comprehensibility (average 3.50 points) also improved significantly compared to the first round. The Content Validity Index (CVI) was 1.0 in all areas, confirming very high validity of the design principles, and the Inter-Rater Agreement (IRA) was also 1.0, indicating high reliability of expert evaluation. All 11 individual design principles recorded an average of 3.67 points or higher.
Additionally, a platform was developed to enable teachers to easily design personalized AI chatbots. The platform was built with a serverless architecture utilizing Cloud Run on Google Cloud Platform (GCP), with Svelte for the frontend and FastAPI for the backend. Teachers can monitor individual student conversations and progress through a real-time dashboard.
External validation was conducted in December to verify the field applicability of the developed design principles. Three sessions of instruction were conducted using teacher-designed personalized chatbots at one middle school in Gyeonggi Province (30 third-year students) and one high school in Gyeongsangbuk Province (30 second-year students), and learner and instructor responses were collected and analyzed. Learners recognized immediate accessibility, individualized explanations, and emotional comfort as strengths, while suggesting improvements such as rendering errors in mathematical formulas, use of concepts beyond grade level, and lack of flexible scaffolding. Instructors positively evaluated the validity and usefulness of the design principles while suggesting supplementary explanations for hyperparameters, functions for generating variant problems, and automation of system prompt improvement. Reflecting these external validation results, the final design principles were derived with enhancements including flexible scaffolding, grade-level concept control, strengthened formula rendering guidelines, and variant problem provision functions.
This study has the following limitations. First, although external validation was conducted at one middle school and one high school, the number and scope of research subjects were limited, presenting limitations for generalizing the design principles. Additional field application studies with diverse school levels, regions, and learner characteristics are needed. Second, this study focused on the mathematical problem-solving domain, and additional research is needed for the possibility of extension to other types of mathematics instruction such as concept learning, inquiry activities, and project-based learning, as well as to other subjects. Third, the external validation process did not employ quasi-experimental designs with control groups or quantitative effect verification through pre-post tests. Learner conversation data and responses were analyzed qualitatively, and quantitative effect verification such as changes in academic achievement resulting from the application of design principles needs to be addressed in follow-up studies.
This study has the following significance. First, it established a theoretical foundation by presenting systematic design principles for personalized mathematical problem-solving instruction using generative AI chatbots. While existing studies remained at the level of utilizing general-purpose AI, this study presented the possibility of teachers directly designing student-personalized AI chatbots and utilizing them as AI teaching assistants through student interaction, along with specific guidelines. Second, by developing comprehensive principles integrating chatbot design and instructional design, it proposed a new instructional design approach for the AI era. Third, through empirical data from pilot testing, theoretical refinement through expert validation, and field verification through external validation, both field applicability and academic validity were simultaneously secured.
목차 (Table of Contents)