This study examined differences between ChatGPT and native English-speaking teachers in erroranalysis and writing assessment of English learner compositions, and explored the relationshipbetween error analysis and assessment outcomes. Twenty-seven col...
This study examined differences between ChatGPT and native English-speaking teachers in erroranalysis and writing assessment of English learner compositions, and explored the relationshipbetween error analysis and assessment outcomes. Twenty-seven college English learnersparticipated in the study, and Pearson correlation, multiple regression analysis, and one-wayANOVA were conducted. The results showed that ChatGPT and the teachers shared a broadlysimilar framework for recognizing learner errors, but differed in the scope and focus of erroridentification. While both displayed comparable tendencies in ‘language-use’ errors, ChatGPTidentified a greater number of errors in ‘content’, ‘organization’, and ‘vocabulary’ based onbroader, macro-level criteria, whereas the teachers identified more errors in ‘writing technique’. Inwriting assessment, no significant differences were found in surface-level linguistic domains;however, a significant difference emerged between ChatGPT and one native English teacher in‘content’ and ‘organization’. Finally, error frequency was not applied as a uniform penalty factorby either ChatGPT and the teachers, suggesting that assessment outcomes depended more on howevaluators interpreted and incorporated errors within their evaluative perspectives. Pedagogicalimplications and directions of further research were discussed.