This study aimed to explore a development procedure for a classroom-based literacy diagnostic assessment for seventh-grade students by integrating generative AI, the EBS Reading Index (ERI), and Retrieval-Augmented Generation (RAG) technology. Drawing...
This study aimed to explore a development procedure for a classroom-based literacy diagnostic assessment for seventh-grade students by integrating generative AI, the EBS Reading Index (ERI), and Retrieval-Augmented Generation (RAG) technology. Drawing on the framework of the EBS Literacy Diagnostic Assessment and the 2022 revised national curriculum, the study first specified passage design conditions (domain, topic, key concepts, text type, length) for five domains—arts, social studies, science, humanities, and integrated content. Based on these specifications, a generative AI model (ChatGPT) was used to produce draft passages and 15 multiple-choice items (three items per passage: literal comprehension, vocabulary, and inference/application). The AI-generated texts and items were then reviewed and revised through a two-step quality control process: factual and consistency checks using a RAG-based tool (NotebookLM) and subsequent human review by the researcher. To adjust text difficulty to the target grade level, ERI scores were calculated by combining quantitative indices derived from the National Institute of Korean Language’s vocabulary level lists and sentence complexity measures with qualitative ratings provided by five in-service Korean language teachers; passages were revised so that their ERI scores fell within the recommended range for seventh grade (7.0–8.5).
A pilot test (n = 27) and a main test (n = 199) were conducted with first-year middle school students in Jeonju, South Korea. Classical test theory analyses were carried out, including item difficulty (p-values), item–total correlations, and Cronbach’s alpha for internal consistency, and the relationship between ERI scores and students’ perceived difficulty (6-point Likert scale) was examined. The final five passages showed ERI scores ranging from 7.0 to 8.1, indicating an appropriate difficulty level for the target grade. In both the pilot and main administrations, the rank order of ERI scores across passages exactly matched the rank order of mean perceived difficulty, suggesting that ERI can serve as a practical indicator for calibrating the difficulty of AI-generated texts. Cronbach’s alpha for the 15-item test was .79 in the pilot study and .64 (standardized α = .65) in the main study, indicating an acceptable level of internal consistency for an exploratory classroom- and school-level diagnostic tool, though not yet sufficient for a fully standardized large-scale assessment.
Overall, the findings suggest that generative AI, when combined with ERI-based difficulty control and RAG-supported factual verification, can function as a supportive tool that enhances teachers’ capacity to develop literacy diagnostic assessments while still requiring professional human judgment at key stages. The study proposes a tentative R&D protocol for AI-assisted literacy test construction in school settings and provides empirical evidence of both its potential and its limitations. Given the constraints of a single school, a single grade level, and a small number of passages and items, further research is needed with more diverse samples, expanded item pools, and more sophisticated analytical approaches such as factor analysis, item response theory, and Rasch modeling.