Geography education aims to foster competencies in various domains, including location knowledge and geographic concepts, as well as geographic imagination, graphicacy, and spatial thinking. Consequently, there is a need to evaluate learning outcomes ...
Geography education aims to foster competencies in various domains, including location knowledge and geographic concepts, as well as geographic imagination, graphicacy, and spatial thinking. Consequently, there is a need to evaluate learning outcomes from multiple perspectives, leading to increased interest in supply-type items. However, the practical application of supply-type items has been limited due to issues such as subjectivity in assessment, excessive workload, and time constraints. To overcome these limitations, automated assessment technology using AI has recently emerged as an alternative. Nevertheless, it has been pointed out that existing automated assessment methods utilizing LLM suffer from hallucinations and a failure to reflect the specific characteristics of individual subjects. Therefore, this study developed an AI automated assessment platform tailored to specific types of supply-type items in geography using the RAG framework and MLLM, and verified the reliability of the platform.
For platform development, Python was adopted as the primary language, and the platform was designed by combining AI models and specific libraries suited to the research purpose. Furthermore, reflecting the characteristics of geography supply-type items, assessment logic was designed to assess both text-based and image-based responses. The RAG framework was applied to enable the LLM to reflect specific geographic curriculum knowledge, and MLLM models (Gemini-2.5-flash and GPT-5-mini) were utilized to perform automated assessment even for responses drawn on outline maps. Finally, the platform was implemented to operate on the web using Streamlit and LangChain.
After the platform construction, a total of three assessment items and assessment rubrics, including both text and image formats, were developed to analyze reliability. For assessment, responses were collected from 85 second-year students at a general high school in Gyeonggi-do, and assessment was performed by both the automated assessment platform and five in-service geography teachers. Intra-rater reliability was verified using descriptive statistics, RM ANOVA, and ICC to check the consistency of the AI models across assessment sessions. Inter-rater reliability was analyzed using descriptive statistics, correlation analysis, and ICC to examine the consistency between automated assessment and teacher assessment.
The analysis results indicated that the developed automated assessment platform generally demonstrated stable consistency across assessment sessions and secured a respectable level of reliability comparable to that of geography teachers. In the analysis of intra-rater reliability, both Gemini and GPT models confirmed a significant level of reliability for text items with clear criteria for correctness as well as those requiring diverse ideas. However, for image-based items, the Gemini model committed errors by grading points to blank responses. When blank responses were filtered out, the Gemini model also secured a reliability level above a certain threshold. In the inter-rater reliability analysis, the GPT model demonstrated a high level of reliability comparable to that of teachers, while the Gemini model showed slightly lower reliability. Nevertheless, the reliability of the average scores compared to teachers for both models reached a 'very good' level, indicating significantly high consistency.
This study holds significance by providing a platform of practical utility for geography education assessment and demonstrating that such a system serves as a valid instrument in the field of evaluation. It is anticipated that this work will catalyze further research on the implementation and assessment of various supply-type items designed to measure learning outcomes in geography.