RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    한국어 학습자 쓰기 능력 진단을 위한 자동 평가 모델 구축 연구 : 과제 독립적 언어 자질을 중심으로

    한글로보기

    https://www.riss.kr/link?id=T17551764

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study aims to develop a machine learning-based automated writing evaluation model using task-independent linguistic features to diagnose the writing proficiency of Korean language learners and to identify the key features contributing to proficiency prediction. A total of 14,992 written texts from the National Institute of Korean Language (NIKL) Korean Learner Corpus were analyzed. Linguistic features were extracted across three domains: lexical complexity, syntactic complexity, and surface features. RandomForest and XGBoost models were trained, and their performance was compared under both 3-level and 6-level classification schemes. Model predictions were then interpreted using SHAP (SHapley Additive exPlanations) analysis.
    Lexical complexity was operationalized through five subcategories: lexical use, lexical diversity, lexical density, lexical sophistication, and collocation use. Syntactic complexity comprised grammatical use, grammatical difficulty, sentence complexity, clausal complexity, and phrasal complexity. Surface features included text length and part-of-speech distribution. Lexical and grammatical difficulty were quantified using the vocabulary and grammar lists from the International Standard Curriculum for Korean Language Education (2017) as reference inventories. Collocation use was measured through N-gram analysis, and phrasal complexity was measured through dependency parsing by computing node depth and the number of dependent nodes.
    Statistical analyses revealed significant differences across proficiency groups, with large effect sizes for key features, including the number of lexical tokens and types per sentence (ω² = .615, .634), the proportion of Level 1 vocabulary types (ω² = .678), and the proportion of Level 1 grammar types (ω² = .697). Correlation analyses further confirmed strong associations between proficiency and the numbers of lexical tokens and types per sentence (ρ = .819, .820), mean sentence length in eojeol units (ρ = .806), and the mean number of dependent nodes (ρ = .805). As proficiency increased, reliance on basic vocabulary decreased while the use of intermediate and advanced vocabulary increased. Adnominal endings showed significant differences across all proficiency groups, functioning as a key indicator of syntactic development.
    XGBoost outperformed RandomForest under both classification schemes, achieving an accuracy of 0.894 in the 3-level classification and 0.770 in the 6-level classification. Although overall accuracy decreased under the finer-grained classification, ROC-AUC values remained above 0.91 across all levels in the 6-level scheme, indicating robust discriminative performance. Classification accuracy was highest for Levels 1 and 2, while greater confusion between adjacent levels was observed in the Level 3–4 range, reflecting the continuous nature of proficiency development at the intermediate stage. These results suggest that the 6-level scheme is suitable for automated assessment, given its direct alignment with TOPIK and the International Standard Curriculum for Korean Language Education.
    SHAP analysis revealed that the number of Level 4 vocabulary types and the proportion of Level 1 vocabulary types were the most influential features for proficiency prediction in both models. At the beginner level, high proportions of Level 1 vocabulary and grammar served as key classification signals. At the intermediate level, increased use of Level 3 grammar and decreased reliance on Level 1 grammar were the primary discriminators. At the advanced level, diverse use of intermediate-to-advanced vocabulary and a lower proportion of sentence-final endings contributed to classification outcomes. These findings demonstrate that SHAP analysis can represent learners’ stage-wise linguistic development in an interpretable form.
    This study makes several contributions to the field. First, it integrates task-independent features—stable across topics and scoring criteria—into a curriculum-aligned framework spanning lexical and syntactic complexity. Second, global and local SHAP analyses address the black-box limitations of tree-based models and yield educationally interpretable rationales, supporting a prototype diagnostic tool for teachers. Third, by covering all six proficiency levels using a large-scale learner corpus, it overcomes the limitations of prior work in terms of data size and proficiency scope.
    번역하기

    This study aims to develop a machine learning-based automated writing evaluation model using task-independent linguistic features to diagnose the writing proficiency of Korean language learners and to identify the key features contributing to proficie...

    This study aims to develop a machine learning-based automated writing evaluation model using task-independent linguistic features to diagnose the writing proficiency of Korean language learners and to identify the key features contributing to proficiency prediction. A total of 14,992 written texts from the National Institute of Korean Language (NIKL) Korean Learner Corpus were analyzed. Linguistic features were extracted across three domains: lexical complexity, syntactic complexity, and surface features. RandomForest and XGBoost models were trained, and their performance was compared under both 3-level and 6-level classification schemes. Model predictions were then interpreted using SHAP (SHapley Additive exPlanations) analysis.
    Lexical complexity was operationalized through five subcategories: lexical use, lexical diversity, lexical density, lexical sophistication, and collocation use. Syntactic complexity comprised grammatical use, grammatical difficulty, sentence complexity, clausal complexity, and phrasal complexity. Surface features included text length and part-of-speech distribution. Lexical and grammatical difficulty were quantified using the vocabulary and grammar lists from the International Standard Curriculum for Korean Language Education (2017) as reference inventories. Collocation use was measured through N-gram analysis, and phrasal complexity was measured through dependency parsing by computing node depth and the number of dependent nodes.
    Statistical analyses revealed significant differences across proficiency groups, with large effect sizes for key features, including the number of lexical tokens and types per sentence (ω² = .615, .634), the proportion of Level 1 vocabulary types (ω² = .678), and the proportion of Level 1 grammar types (ω² = .697). Correlation analyses further confirmed strong associations between proficiency and the numbers of lexical tokens and types per sentence (ρ = .819, .820), mean sentence length in eojeol units (ρ = .806), and the mean number of dependent nodes (ρ = .805). As proficiency increased, reliance on basic vocabulary decreased while the use of intermediate and advanced vocabulary increased. Adnominal endings showed significant differences across all proficiency groups, functioning as a key indicator of syntactic development.
    XGBoost outperformed RandomForest under both classification schemes, achieving an accuracy of 0.894 in the 3-level classification and 0.770 in the 6-level classification. Although overall accuracy decreased under the finer-grained classification, ROC-AUC values remained above 0.91 across all levels in the 6-level scheme, indicating robust discriminative performance. Classification accuracy was highest for Levels 1 and 2, while greater confusion between adjacent levels was observed in the Level 3–4 range, reflecting the continuous nature of proficiency development at the intermediate stage. These results suggest that the 6-level scheme is suitable for automated assessment, given its direct alignment with TOPIK and the International Standard Curriculum for Korean Language Education.
    SHAP analysis revealed that the number of Level 4 vocabulary types and the proportion of Level 1 vocabulary types were the most influential features for proficiency prediction in both models. At the beginner level, high proportions of Level 1 vocabulary and grammar served as key classification signals. At the intermediate level, increased use of Level 3 grammar and decreased reliance on Level 1 grammar were the primary discriminators. At the advanced level, diverse use of intermediate-to-advanced vocabulary and a lower proportion of sentence-final endings contributed to classification outcomes. These findings demonstrate that SHAP analysis can represent learners’ stage-wise linguistic development in an interpretable form.
    This study makes several contributions to the field. First, it integrates task-independent features—stable across topics and scoring criteria—into a curriculum-aligned framework spanning lexical and syntactic complexity. Second, global and local SHAP analyses address the black-box limitations of tree-based models and yield educationally interpretable rationales, supporting a prototype diagnostic tool for teachers. Third, by covering all six proficiency levels using a large-scale learner corpus, it overcomes the limitations of prior work in terms of data size and proficiency scope.

    더보기

    목차 (Table of Contents)

    • 1. 서론 1
    • 1.1 연구 필요성 및 목적 1
    • 1.2 연구 문제 7
    • 1.3 선행 연구 9
    • 1.3.1 언어 자질 관련 연구 9
    • 1. 서론 1
    • 1.1 연구 필요성 및 목적 1
    • 1.2 연구 문제 7
    • 1.3 선행 연구 9
    • 1.3.1 언어 자질 관련 연구 9
    • 1.3.2 자동 채점 관련 연구 18
    • 2. 이론적 배경 25
    • 2.1 쓰기 능력 25
    • 2.1.1 쓰기 능력의 개념과 구성 요소 25
    • 2.1.2 등급별 쓰기 평가 기준 29
    • 2.2 쓰기 능력 측정을 위한 언어 자질 33
    • 2.2.1 어휘 복잡도 33
    • 2.2.2 통사 복잡도 43
    • 2.3 해석 가능한 머신러닝 50
    • 2.3.1 트리 기반 모델 52
    • 2.3.2 모델 해석 기법 56
    • 3. 연구 방법 60
    • 3.1 코퍼스 선정 61
    • 3.2 데이터 전처리 65
    • 3.2.1 참조 어휘 목록 65
    • 3.2.2 참조 문법 목록 67
    • 3.3 언어 자질 선정 71
    • 3.3.1 어휘 복잡도 관련 자질 71
    • 3.3.2 통사 복잡도 관련 자질 77
    • 3.3.3 표층 자질 82
    • 3.4 예측 모델 구축 86
    • 3.4.1 데이터 분할 86
    • 3.4.2 모델 학습 88
    • 3.4.3 모델 평가 지표 90
    • 4. 등급별 언어 발달 양상 분석 93
    • 4.1 자질별 기술 통계 분석 93
    • 4.1.1 어휘 복잡도 관련 자질 93
    • 4.1.2 통사 복잡도 관련 자질 106
    • 4.1.3 표층 자질 115
    • 4.2 ANOVA 및 상관관계 분석 118
    • 4.2.1 어휘 복잡도 관련 자질 118
    • 4.2.2 통사 복잡도 관련 자질 127
    • 4.2.3 표층 자질 134
    • 4.2.4 분석 결과 종합 137
    • 5. 모델 평가 및 해석 141
    • 5.1 모델 성능 평가 141
    • 5.1.1 3등급 분류 체계 141
    • 5.1.2 6등급 분류 체계 146
    • 5.2 모델 해석 결과 151
    • 5.2.1 자질 중요도 분석 151
    • 5.2.2 SHAP 분석 158
    • 5.3 모델 기반 사례 분석 179
    • 5.3.1 대표 사례 분석 179
    • 6. 교사 보조 도구로서 모델 활용 방안 196
    • 6.1 신뢰도 기준 설정 198
    • 6.2 자동 평가 프로그램 구현 201
    • 7. 결론 및 제언 208
    • 참고문헌 213
    • ABSTRACT 234
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼