RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    지리과 서답형 문항 AI 자동 채점 플랫폼 개발 = Development of an AI-Based Automated Assessment Platform for Supply-Type Items in Geography Education

    한글로보기

    https://www.riss.kr/link?id=T17451848

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Geography education aims to foster competencies in various domains, including location knowledge and geographic concepts, as well as geographic imagination, graphicacy, and spatial thinking. Consequently, there is a need to evaluate learning outcomes from multiple perspectives, leading to increased interest in supply-type items. However, the practical application of supply-type items has been limited due to issues such as subjectivity in assessment, excessive workload, and time constraints. To overcome these limitations, automated assessment technology using AI has recently emerged as an alternative. Nevertheless, it has been pointed out that existing automated assessment methods utilizing LLM suffer from hallucinations and a failure to reflect the specific characteristics of individual subjects. Therefore, this study developed an AI automated assessment platform tailored to specific types of supply-type items in geography using the RAG framework and MLLM, and verified the reliability of the platform.
    For platform development, Python was adopted as the primary language, and the platform was designed by combining AI models and specific libraries suited to the research purpose. Furthermore, reflecting the characteristics of geography supply-type items, assessment logic was designed to assess both text-based and image-based responses. The RAG framework was applied to enable the LLM to reflect specific geographic curriculum knowledge, and MLLM models (Gemini-2.5-flash and GPT-5-mini) were utilized to perform automated assessment even for responses drawn on outline maps. Finally, the platform was implemented to operate on the web using Streamlit and LangChain.
    After the platform construction, a total of three assessment items and assessment rubrics, including both text and image formats, were developed to analyze reliability. For assessment, responses were collected from 85 second-year students at a general high school in Gyeonggi-do, and assessment was performed by both the automated assessment platform and five in-service geography teachers. Intra-rater reliability was verified using descriptive statistics, RM ANOVA, and ICC to check the consistency of the AI models across assessment sessions. Inter-rater reliability was analyzed using descriptive statistics, correlation analysis, and ICC to examine the consistency between automated assessment and teacher assessment.
    The analysis results indicated that the developed automated assessment platform generally demonstrated stable consistency across assessment sessions and secured a respectable level of reliability comparable to that of geography teachers. In the analysis of intra-rater reliability, both Gemini and GPT models confirmed a significant level of reliability for text items with clear criteria for correctness as well as those requiring diverse ideas. However, for image-based items, the Gemini model committed errors by grading points to blank responses. When blank responses were filtered out, the Gemini model also secured a reliability level above a certain threshold. In the inter-rater reliability analysis, the GPT model demonstrated a high level of reliability comparable to that of teachers, while the Gemini model showed slightly lower reliability. Nevertheless, the reliability of the average scores compared to teachers for both models reached a 'very good' level, indicating significantly high consistency.
    This study holds significance by providing a platform of practical utility for geography education assessment and demonstrating that such a system serves as a valid instrument in the field of evaluation. It is anticipated that this work will catalyze further research on the implementation and assessment of various supply-type items designed to measure learning outcomes in geography.
    번역하기

    Geography education aims to foster competencies in various domains, including location knowledge and geographic concepts, as well as geographic imagination, graphicacy, and spatial thinking. Consequently, there is a need to evaluate learning outcomes ...

    Geography education aims to foster competencies in various domains, including location knowledge and geographic concepts, as well as geographic imagination, graphicacy, and spatial thinking. Consequently, there is a need to evaluate learning outcomes from multiple perspectives, leading to increased interest in supply-type items. However, the practical application of supply-type items has been limited due to issues such as subjectivity in assessment, excessive workload, and time constraints. To overcome these limitations, automated assessment technology using AI has recently emerged as an alternative. Nevertheless, it has been pointed out that existing automated assessment methods utilizing LLM suffer from hallucinations and a failure to reflect the specific characteristics of individual subjects. Therefore, this study developed an AI automated assessment platform tailored to specific types of supply-type items in geography using the RAG framework and MLLM, and verified the reliability of the platform.
    For platform development, Python was adopted as the primary language, and the platform was designed by combining AI models and specific libraries suited to the research purpose. Furthermore, reflecting the characteristics of geography supply-type items, assessment logic was designed to assess both text-based and image-based responses. The RAG framework was applied to enable the LLM to reflect specific geographic curriculum knowledge, and MLLM models (Gemini-2.5-flash and GPT-5-mini) were utilized to perform automated assessment even for responses drawn on outline maps. Finally, the platform was implemented to operate on the web using Streamlit and LangChain.
    After the platform construction, a total of three assessment items and assessment rubrics, including both text and image formats, were developed to analyze reliability. For assessment, responses were collected from 85 second-year students at a general high school in Gyeonggi-do, and assessment was performed by both the automated assessment platform and five in-service geography teachers. Intra-rater reliability was verified using descriptive statistics, RM ANOVA, and ICC to check the consistency of the AI models across assessment sessions. Inter-rater reliability was analyzed using descriptive statistics, correlation analysis, and ICC to examine the consistency between automated assessment and teacher assessment.
    The analysis results indicated that the developed automated assessment platform generally demonstrated stable consistency across assessment sessions and secured a respectable level of reliability comparable to that of geography teachers. In the analysis of intra-rater reliability, both Gemini and GPT models confirmed a significant level of reliability for text items with clear criteria for correctness as well as those requiring diverse ideas. However, for image-based items, the Gemini model committed errors by grading points to blank responses. When blank responses were filtered out, the Gemini model also secured a reliability level above a certain threshold. In the inter-rater reliability analysis, the GPT model demonstrated a high level of reliability comparable to that of teachers, while the Gemini model showed slightly lower reliability. Nevertheless, the reliability of the average scores compared to teachers for both models reached a 'very good' level, indicating significantly high consistency.
    This study holds significance by providing a platform of practical utility for geography education assessment and demonstrating that such a system serves as a valid instrument in the field of evaluation. It is anticipated that this work will catalyze further research on the implementation and assessment of various supply-type items designed to measure learning outcomes in geography.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    지리교육은 위치 지식, 지리적 개념과 더불어 지리적 상상력, 도해력, 공간적 사고력 등 다양한 영역에서의 역량 함양을 목표로 한다. 따라서 다방면에서 학습 성과를 평가할 필요가 있고, 이러한 맥락에서 서답형 문항에 대한 관심이 커지고 있다. 그러나 서답형 문항은 채점 주관성 시비, 과도한 업무 부담, 시간적 제약 등으로 인해 실질적으로 활용하는 데 제한이 있었다. 이러한 한계를 극복하기 위해 최근 AI를 활용한 자동 채점 기술이 대안으로 부상하고 있으나, 기존 LLM 활용 자동 채점은 환각 현상과 개별 교과의 특수성을 반영하지 못하는 문제점이 지적되었다. 이에 본 연구에서는 RAG 프레임워크와 MLLM을 활용하여 지리과 서답형 문항 AI 자동 채점 플랫폼을 개발하고, 해당 플랫폼의 신뢰도를 검증하였다.
    플랫폼 개발을 위해 Python을 주 언어로 채택하였으며, 연구 목적에 맞는 AI 모델과 세부 라이브러리를 조합하여 플랫폼을 고안하였다. 또한 지리과 서답형 문항의 특성을 반영하여 텍스트와 이미지 형태 답안을 모두 채점할 수 있도록 채점 로직을 설계하였다. LLM이 지리 교과 지식을 반영할 수 있도록 RAG 프레임워크를 적용하였으며, MLLM 모델(Gemini-2.5-flash 및 GPT-5-mini)을 활용하여 백지도에 작성한 답안에 대해서도 자동 채점을 수행할 수 있도록 하였다. 최종적으로 Streamlit과 LangChain을 활용하여 웹에서 작동되도록 구현하였다.
    플랫폼 구축 이후, 신뢰도를 분석하기 위해 텍스트와 이미지 형태의 답안을 포함하는 총 3개의 평가 문항 및 채점 루브릭을 개발하였다. 채점을 위해 경기도 소재 일반계 고등학교 2학년 85명의 답안을 수집하였으며, 자동 채점 플랫폼과 현직 지리 교사 5명이 채점을 수행하였다. 평가자 내 신뢰도는 기술 통계, RM ANOVA, ICC를 통해 AI 모델의 채점 회차별 일관성을 확인하였고, 평가자 간 신뢰도는 기술 통계, 상관분석, ICC를 통해 자동 채점과 지리 교사 채점 간 일관성을 분석하였다.
    분석 결과, 개발한 자동 채점 플랫폼은 전반적으로 안정적인 채점 회차별 일관성을 보였으며, 지리 교사와 비교하여도 준수한 신뢰도를 확보하였다. 평가자 내 신뢰도 분석에서 Gemini와 GPT 모델 모두 정오답 판단 기준이 명확한 텍스트 문항과 다양한 아이디어를 묻는 텍스트 문항에 대해서 상당한 수준의 신뢰도를 확인하였다. 그러나 이미지 형태의 문항에서 Gemini 모델은 공백 답안에 대해 점수를 부여하는 오류를 범하였다. 공백인 답안을 정제한 경우, Gemini 모델도 일정 수준 이상의 신뢰도를 확보하였다. 평가자 간 신뢰도 분석에서는 GPT 모델은 교사와 비교하여도 대등한 수준의 높은 신뢰도를 보였으나, Gemini 모델은 약간 낮은 신뢰도를 보였다. 교사와의 평균 점수의 신뢰도에서는 두 모델 모두 ‘매우 좋음’ 수준의 상당히 높은 신뢰도가 나타났다.
    본 연구는 지리교육 평가에 실질적으로 활용할 수 있는 플랫폼을 제공하고, 해당 플랫폼이 평가 현장에서 의미 있는 역할을 수행할 수 있음을 밝혔다는 점에서 의의가 있다. 향후 지리 교과의 학습 성과를 측정할 수 있는 다양한 서답형 문항의 도입과 채점에 대한 연구가 이어지기를 기대한다.
    번역하기

    지리교육은 위치 지식, 지리적 개념과 더불어 지리적 상상력, 도해력, 공간적 사고력 등 다양한 영역에서의 역량 함양을 목표로 한다. 따라서 다방면에서 학습 성과를 평가할 필요가 있고, ...

    지리교육은 위치 지식, 지리적 개념과 더불어 지리적 상상력, 도해력, 공간적 사고력 등 다양한 영역에서의 역량 함양을 목표로 한다. 따라서 다방면에서 학습 성과를 평가할 필요가 있고, 이러한 맥락에서 서답형 문항에 대한 관심이 커지고 있다. 그러나 서답형 문항은 채점 주관성 시비, 과도한 업무 부담, 시간적 제약 등으로 인해 실질적으로 활용하는 데 제한이 있었다. 이러한 한계를 극복하기 위해 최근 AI를 활용한 자동 채점 기술이 대안으로 부상하고 있으나, 기존 LLM 활용 자동 채점은 환각 현상과 개별 교과의 특수성을 반영하지 못하는 문제점이 지적되었다. 이에 본 연구에서는 RAG 프레임워크와 MLLM을 활용하여 지리과 서답형 문항 AI 자동 채점 플랫폼을 개발하고, 해당 플랫폼의 신뢰도를 검증하였다.
    플랫폼 개발을 위해 Python을 주 언어로 채택하였으며, 연구 목적에 맞는 AI 모델과 세부 라이브러리를 조합하여 플랫폼을 고안하였다. 또한 지리과 서답형 문항의 특성을 반영하여 텍스트와 이미지 형태 답안을 모두 채점할 수 있도록 채점 로직을 설계하였다. LLM이 지리 교과 지식을 반영할 수 있도록 RAG 프레임워크를 적용하였으며, MLLM 모델(Gemini-2.5-flash 및 GPT-5-mini)을 활용하여 백지도에 작성한 답안에 대해서도 자동 채점을 수행할 수 있도록 하였다. 최종적으로 Streamlit과 LangChain을 활용하여 웹에서 작동되도록 구현하였다.
    플랫폼 구축 이후, 신뢰도를 분석하기 위해 텍스트와 이미지 형태의 답안을 포함하는 총 3개의 평가 문항 및 채점 루브릭을 개발하였다. 채점을 위해 경기도 소재 일반계 고등학교 2학년 85명의 답안을 수집하였으며, 자동 채점 플랫폼과 현직 지리 교사 5명이 채점을 수행하였다. 평가자 내 신뢰도는 기술 통계, RM ANOVA, ICC를 통해 AI 모델의 채점 회차별 일관성을 확인하였고, 평가자 간 신뢰도는 기술 통계, 상관분석, ICC를 통해 자동 채점과 지리 교사 채점 간 일관성을 분석하였다.
    분석 결과, 개발한 자동 채점 플랫폼은 전반적으로 안정적인 채점 회차별 일관성을 보였으며, 지리 교사와 비교하여도 준수한 신뢰도를 확보하였다. 평가자 내 신뢰도 분석에서 Gemini와 GPT 모델 모두 정오답 판단 기준이 명확한 텍스트 문항과 다양한 아이디어를 묻는 텍스트 문항에 대해서 상당한 수준의 신뢰도를 확인하였다. 그러나 이미지 형태의 문항에서 Gemini 모델은 공백 답안에 대해 점수를 부여하는 오류를 범하였다. 공백인 답안을 정제한 경우, Gemini 모델도 일정 수준 이상의 신뢰도를 확보하였다. 평가자 간 신뢰도 분석에서는 GPT 모델은 교사와 비교하여도 대등한 수준의 높은 신뢰도를 보였으나, Gemini 모델은 약간 낮은 신뢰도를 보였다. 교사와의 평균 점수의 신뢰도에서는 두 모델 모두 ‘매우 좋음’ 수준의 상당히 높은 신뢰도가 나타났다.
    본 연구는 지리교육 평가에 실질적으로 활용할 수 있는 플랫폼을 제공하고, 해당 플랫폼이 평가 현장에서 의미 있는 역할을 수행할 수 있음을 밝혔다는 점에서 의의가 있다. 향후 지리 교과의 학습 성과를 측정할 수 있는 다양한 서답형 문항의 도입과 채점에 대한 연구가 이어지기를 기대한다.

    더보기

    목차 (Table of Contents)

    • 제 1 장 서론 1
    • 제 1 절 연구 필요성과 목적 1
    • 제 2 절 연구 문제 4
    • 제 2 장 이론적 배경 5
    • 제 1 장 서론 1
    • 제 1 절 연구 필요성과 목적 1
    • 제 2 절 연구 문제 4
    • 제 2 장 이론적 배경 5
    • 제 1 절 지리과 서답형 문항 5
    • 1. 응답 방식에 따른 평가 문항 유형 5
    • 2. 서답형 문항 유형 8
    • 3. 지리과 서답형 문항 유형 10
    • 제 2 절 AI 활용 자동 채점 12
    • 1. 자동 채점 12
    • 2. LLM 활용 자동 채점 14
    • 3. MLLM 활용 자동 채점 17
    • 제 3 장 연구 방법 20
    • 제 1 절 연구 절차 20
    • 제 2 절 연구 참여자 22
    • 1. 고등학생 참여자 22
    • 2. 지리 교사 참여자 23
    • 제 3 절 검사 도구 24
    • 1. 연구용 지리과 서답형 평가 문항 24
    • 1) 평가 문항 1번: 제한형 / 지도 자료 분석형 / 이해 요구형 25
    • 2) 평가 문항 2번: 확장형 / 글 자료 제시형 / 적용 요구형 26
    • 3) 평가 문항 3번: 단답형 / 지도 자료 수행형 / 지식 요구형 29
    • 2. 지리과 서답형 평가 문항별 채점 루브릭 31
    • 1) 평가 문항 1번 채점 루브릭 31
    • 2) 평가 문항 2번 채점 루브릭 33
    • 3) 평가 문항 3번 채점 루브릭 34
    • 제 4 절 플랫폼 개발 도구 36
    • 1. 플랫폼 개발 환경 36
    • 2. 자동 채점 수행 AI 모델 선정 36
    • 3. RAG 프레임워크 설계 38
    • 제 5 절 분석 방법 40
    • 1. 자동 채점 플랫폼의 평가자 내 신뢰도 40
    • 2. 자동 채점 플랫폼의 평가자 간 신뢰도 43
    • 제 4 장 연구 결과 47
    • 제 1 절 지리과 서답형 문항 AI 자동 채점 플랫폼 개발 47
    • 1. 지리과 서답형 문항 채점 로직 및 프롬프트 설계 47
    • 1) 텍스트 문항 채점 로직 및 프롬프트 설계 47
    • 2) 백지도 문항 채점 로직 및 프롬프트 설계 52
    • 2. 플랫폼 아키텍처 및 주요 기능 설계 57
    • 1) 채점 시스템 설정 페이지 57
    • 2) 평가 루브릭 입력 페이지 61
    • 3) 채점 실행 페이지 63
    • 4) 채점 결과 출력 페이지 64
    • 제 2 절 자동 채점 플랫폼의 평가자 내 신뢰도 분석 결과 66
    • 1. 문항 1번에 대한 평가자 내 신뢰도 분석 결과 66
    • 2. 문항 2번에 대한 평가자 내 신뢰도 분석 결과 69
    • 3. 문항 3번에 대한 평가자 내 신뢰도 분석 결과 71
    • 4. 자동 채점 플랫폼의 평가자 내 신뢰도 분석에 대한 논의 75
    • 제 3 절 자동 채점 플랫폼과 지리 교사의 평가자 간 신뢰도 분석 78
    • 1. 문항 1번에 대한 평가자 간 신뢰도 분석 결과 78
    • 2. 문항 2번에 대한 평가자 간 신뢰도 분석 결과 82
    • 3. 문항 3번에 대한 평가자 간 신뢰도 분석 결과 85
    • 4. 자동 채점 플랫폼의 평가자 간 신뢰도 분석에 대한 논의 89
    • 제 5 장 결론 96
    • 제 1 절 요약 96
    • 제 2 절 함의 98
    • 제 3 절 한계 및 제언 100
    • 참고문헌 103
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼