An Automatic Assessment Model for Mathematical Expression Answers Combining Contextual Similarity and Formula Equivalence Verification Mun Hee Suk Advisor : Prof. Shin JuHyun, Ph.D. Department of Software Convergence Engineering Graduate School of Fut...
An Automatic Assessment Model for Mathematical Expression Answers Combining Contextual Similarity and Formula Equivalence Verification Mun Hee Suk Advisor : Prof. Shin JuHyun, Ph.D. Department of Software Convergence Engineering Graduate School of Future Human Resources Convergence National technical qualification practical examinations that include descriptive mathematical expression problems can assess not only examinees’ final answ ers but also their calculation processes. However, manual scoring requires co nsiderable time and manpower, and maintaining consistency among raters re mains a challenge. This study proposes an automatic assessment model for mathematical expr ession answers by combining contextual similarity and formula equivalence ve rification. In the first stage, a mathematics-specialized pretrained language m odel is used to embed both model answers and examinee responses, and the cosine similarity between the two embeddings is calculated. Based on this si milarity score, responses are preliminarily classified as correct or incorrect. In the second stage, formula equivalence verification is selectively applied to res ponses that are difficult to determine confidently through the first-stage class ification. This verification process consists of formula normalization, exact ma tching after normalization, symbolic equivalence checking for expressions and equations, and numerical approximation checking. The proposed model does not apply formula equivalence verification to all responses. Instead, it identifies a verification interval around the first-stage cl assification threshold using validation data and applies the second-stage verif ication only to responses within this boundary region. The verification interval is selected through validation data–based performance comparison, and the s elected decision rule is then applied to the test dataset. The experimental data are drawn from the “Mathematics Automatic Solution Data” provided by AI-Hub. The experimental results show that the proposed model consistently outperforms the first-stage MathBERT classifier on both va lidation and test datasets in terms of Accuracy, Balanced Accuracy, and Macro F1-score. In comparison with baseline models, including TF-IDF, KoSBERT, a nd GPT-based approaches, the proposed model achieves the best performan ce across these three evaluation metrics. This study is meaningful in that it demonstrates the feasibility of automatic assessment for mathematical expression answers in an environment where ha ndwritten responses are converted into text through OCR. In particular, the re sults show that combining contextual similarity with formula equivalence verific ation improves both overall accuracy and balanced classification performance. The proposed model provides a practical foundation for applying automatic gr ading to national technical qualification practical examinations. Its performance is expected to improve further as computer-based testing environments and structured mathematical input tools, such as equation editors, become more widely available.