Descriptive assessments are an important evaluation method capable of assessing students' diverse competencies, such as creativity and higher-order thinking skills. However, their implementation in schools has faced significant challenges due to the s...
Descriptive assessments are an important evaluation method capable of assessing students' diverse competencies, such as creativity and higher-order thinking skills. However, their implementation in schools has faced significant challenges due to the substantial time and cost required for grading, coupled with difficulties in ensuring inter-rater reliability. Recently emerged artificial intelligence technology has shown new potential to overcome these limitations. Therefore, this study analyzed the effectiveness of an automated grading program utilizing the latest AI technologies—the natural language generation model (GPT model) and the natural language understanding model (BERT model)—as a method to reduce teachers' burden in expanding essay-type assessments in schools and to enhance assessment reliability. This could contribute to increasing the practical applicability of automated grading.
Specifically, descriptive questions on the concept of ‘demand’ from the middle school social studies and economics unit were developed by categorizing them according to the content framework categories of the 2022 revised curriculum: ‘Knowledge and Understanding’, ‘Process and Function’, and ‘Values and Attitudes’. The automatic grading performance and characteristics of the natural language generation model and natural language understanding model were then compared for each question type. For this study, a total of 900 answer data points collected from 300 third-year middle school students at three middle schools in Seoul were utilized. Student answers were graded using an automatic grading program, and the results were compared and analyzed against teacher grading. The natural language generation model employed an automatic grading program based on Google Sheets, while the natural language understanding model utilized an automatic grading program built on Python within Google Colab for the research.
The main findings of this study are as follows. First, we derived the optimal prompt type and automatic scoring method for each AI model. Natural language generation models showed the highest performance with the ‘few-shot’ type, which provides rubrics, example answers, and specific scoring cases. Natural language understanding models demonstrated high performance with the ‘classification-based’ method, which predicts score intervals by learning patterns in correct answer data. Second, the most effective AI model differed depending on the content domain category of the item. For items in the ‘Knowledge/Understanding’ content framework category, where correct answers are clear, both models showed a very high level of agreement with teacher grading (QWK ≥ 0.9). For items in the ‘Process/Function’ content framework category, requiring logical thinking, the natural language understanding model proved effective. For items in the ‘Value/Attitude’ content framework category, involving the judgment of diverse values in the affective domain, the natural language generation model was the effective model. Third, differences in grading tendencies emerged between models. The natural language generation model exhibited a strict tendency toward ‘undergrading’ due to the influence of hallucination prevention training. Conversely, the natural language understanding model showed a relatively lenient tendency toward ‘overgrading,’ reacting sensitively to whether keywords were present in student responses. These findings suggest that AI-powered automated grading can be an effective tool to assist teachers in grading within school settings. However, rather than applying a single AI model indiscriminately, it is necessary to selectively apply appropriate models and optimal strategies by considering the nature of the item according to the content framework category and the purpose of the assessment (e.g., strict selection or generous feedback).
By proposing optimal AI automatic grading approaches for each content system category, this study is expected to contribute to enhancing the quality of essay-type assessments by reducing the burden teachers feel in schools. It also holds significance in providing foundational data for building AI-based assessment systems.
However, the GPT and BERT models used in this study are based on the most recent models available at this time. The possibility that more effective strategies and methods may emerge in new assessment contexts when newer models are released in the future suggests the need for ongoing research in automated grading. Furthermore, since some scoring models require additional effort for teachers to use directly in school settings, further support will be needed to devise ways to utilize them in more convenient and user-friendly ways.