This study aims to implement an automated and objective model for assessing grammatical competence in text generation, focusing on learners’ developmental patterns based on linguistic features across the lexical, grammatical morphemic, syntactic, an...
This study aims to implement an automated and objective model for assessing grammatical competence in text generation, focusing on learners’ developmental patterns based on linguistic features across the lexical, grammatical morphemic, syntactic, and textual dimensions.
For this purpose, grammatical competence was conceptualized in text generation as “the competence to select appropriate vocabulary, assign syntactic and grammatical functions to the vocabulary using grammatical morphemes such as particles and endings, syntactically combine these elements to form sentences, and connect these sentences to generate coherent texts for expressing intended content”. This concept was operationalized into measurable linguistic indices across the four dimensions mentioned above.
Subsequently, to explore the developmental patterns of elementary and secondary students across the lexical, grammatical morphemic, syntactic, and textual dimensions, linguistic features for assessing grammatical competence were selected by analyzing a large-scale corpus of learner compositions. Based on these features, assessment models using XGBoost and CatBoost were implemented, and the significance of the linguistic features within the models was interpreted. Furthermore, the educational implications of the model implementation were also examined from both procedural and consequential perspectives.
Firstly, this study is grounded in the linguistic view that language must be understood as actual linguistic performance by language users, and thus linguistic competence should be assessed based on concrete linguistic performance. A macro-structure of grammatical competence was established, comprising grammatical knowledge, linguistic intuition, grammatical inquiry competence, and grammatical application competence. This framework clarifies the position of grammatical competence in text generation.
Consequently, this study focuses on the grammatical competence actualized through the linguistic performance of text generation, with a particular emphasis on expression generation. As the process of converting intended meaning into linguistic expression involves lexical, grammatical morphemic, syntactic, and textual dimensions, this study posits that grammatical competence in text generation is a comprehensive competence encompassing all these facets.
Subsequently, to operationalize this competence in an empirical and experiential form, the study explored linguistic indices capable of representing the competences within each dimension. As a result, for the lexical dimension, lexical diversity and vocabulary grade were selected as representative linguistic indices. For the grammatical morphemic dimension, a grammatical morpheme activation index and a grammatical morpheme diversity index were chosen. For the syntactic dimension, syntactic complexity was selected, and for the textual dimension, a coherence score based on a large language model and usage patterns of various cohesion devices were selected.
Based on these selected linguistic indices, this study analyzed a corpus of compositions written by elementary and secondary school students to explore their developmental patterns across the lexical, grammatical morphemic, syntactic, and textual dimensions. The findings revealed that learner development occurs complexly across all four dimensions from the 4th grade of elementary school to the 3rd grade of high school. In particular, grammatical competence was found to develop intensively during the transitions from elementary to middle school and from middle to high school.
Based on the above discussions, 21 linguistic features were selected for assessing grammatical competence, and these were used as features to implement an assessment model. The model, based on XGBoost and CatBoost, was finalized through processes including data splitting and construction for training, validation, and testing sets, as well as hyperparameter optimization using Bayesian optimization. The final model demonstrated an error of RMSE 1.3465, MAE 1.02 grade levels, and MAPE 15.33%. It was confirmed that cases with a prediction error of less than 0.5 grade levels accounted for 33.18% of the total, and those with less than 1.5 grade levels accounted for 77.17%. Furthermore, an interpretation of the model's features using the SHAP technique revealed that key linguistic features included: content word HD-D, content word MATTR, and the proportion of grade 4-5 vocabulary (lexical); embedded structure score and modificational complexity score (syntactic); overall grammatical morpheme diversity index and the usage rate of noun-deriving suffixes (grammatical morphemic); and the frequency of overall cohesion devices, demonstrative/anaphoric cohesion devices, and connective endings (textual). Accordingly, the significance of each linguistic feature in the assessment of grammatical competence was explored based on its SHAP value.
The assessment model implemented through these discussions holds educational value not only as a final product but also in its implementation process. From a procedural perspective, its educational implications include “providing a logical framework for a grammar curriculum based on learner development” and “presenting a methodology for exploring learners' lexico-grammatical resources and choice systems”. From a consequential perspective, it “secures objectivity and specificity in assessment through the grammatical competence assessment model”.
However, this study's model is limited as it focuses solely on quantitative indices and expression generation, while neglecting the assessment of qualitative indices and the aspect of meaning creation. Moreover, although it comprehensively examined linguistic features across the four dimensions, the grammatical competence involved in text generation cannot be fully elucidated by the features explored in this study alone. These limitations highlight the necessity for future research. It is anticipated that if future studies develop qualitative assessment models capable of assessing aspects like the appropriateness of expression, or writing competence models focused on meaning creation, a significant step could be taken toward automated and objective assessment through integration with this work.