Automated scoring is a research field that facilitates constructivist educational evaluation by applying recently developed artificial intelligence technology to learner evaluation. As mathematics is an essential subject of the K-12 curriculum, resear...
Automated scoring is a research field that facilitates constructivist educational evaluation by applying recently developed artificial intelligence technology to learner evaluation. As mathematics is an essential subject of the K-12 curriculum, research on automated scoring of constructed mathematical responses is of great importance. However, unlike research on automated scoring in other subjects’ responses, studies for automated scoring of mathematical responses still require further exploration. In this study, we propose classification criteria for mathematical constructed responses including response types and input types, identified 21 studies from 15 academic journals registered in SCOPUS with systematic literature review, and investigated them in detail.
As a result, Research on automated scoring of mathematical constructed responses has gradually increased in the 2010s, and mainly focusing on algebra and functions, number and operations at the secondary school level. With the advancement of artificial intelligence technology, automated scoring models have also evolved from early rule-based and statistical-based models to machine learning-, deep learning-, and large language models. Furthermore, early automated scoring studies which focused solely on single-modal and digital formatted answers are progressing toward automated scoring of multimodal and handwritten answers. Both holistic and analytic scoring methods are used, and stepwise rubric-based approaches have been identified as more suitable for automated scoring in mathematics.
Furthermore, this study identified that current research mainly emphasizes the technical aspects of scoring systems and their model performance. However, the educational validity of automated scoring is ultimately more critical than its technical accuracy and performance itself. Therefore, future research should extend beyond algorithmic development to address practical applications and pedagogical effectiveness in real classroom contexts.
Lastly, this study suggests practical automated scoring scenarios using sample responses devised by the researcher. Different scoring scenarios depending on the method of response collection during instruction are outlined. Examples of recent applications using recent language models are also introduced. These exploraions provide insight into the potential of automated grading for multimodal responses and offer implications for future implementation in education settings.