RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    생성형 AI와 ERI 지수를 활용한 문해력 진단평가 지문 및 문항 제작 방안에 대한 탐색적 연구 = Exploring a Generative AI and ERI(EBS Reading Index) based Apporach to Developing Reading Texts and Test Items for a Literacy Diagnostic Assessment

    한글로보기

    https://www.riss.kr/link?id=T17370187

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study aimed to explore a development procedure for a classroom-based literacy diagnostic assessment for seventh-grade students by integrating generative AI, the EBS Reading Index (ERI), and Retrieval-Augmented Generation (RAG) technology. Drawing on the framework of the EBS Literacy Diagnostic Assessment and the 2022 revised national curriculum, the study first specified passage design conditions (domain, topic, key concepts, text type, length) for five domains—arts, social studies, science, humanities, and integrated content. Based on these specifications, a generative AI model (ChatGPT) was used to produce draft passages and 15 multiple-choice items (three items per passage: literal comprehension, vocabulary, and inference/application). The AI-generated texts and items were then reviewed and revised through a two-step quality control process: factual and consistency checks using a RAG-based tool (NotebookLM) and subsequent human review by the researcher. To adjust text difficulty to the target grade level, ERI scores were calculated by combining quantitative indices derived from the National Institute of Korean Language’s vocabulary level lists and sentence complexity measures with qualitative ratings provided by five in-service Korean language teachers; passages were revised so that their ERI scores fell within the recommended range for seventh grade (7.0–8.5).
    A pilot test (n = 27) and a main test (n = 199) were conducted with first-year middle school students in Jeonju, South Korea. Classical test theory analyses were carried out, including item difficulty (p-values), item–total correlations, and Cronbach’s alpha for internal consistency, and the relationship between ERI scores and students’ perceived difficulty (6-point Likert scale) was examined. The final five passages showed ERI scores ranging from 7.0 to 8.1, indicating an appropriate difficulty level for the target grade. In both the pilot and main administrations, the rank order of ERI scores across passages exactly matched the rank order of mean perceived difficulty, suggesting that ERI can serve as a practical indicator for calibrating the difficulty of AI-generated texts. Cronbach’s alpha for the 15-item test was .79 in the pilot study and .64 (standardized α = .65) in the main study, indicating an acceptable level of internal consistency for an exploratory classroom- and school-level diagnostic tool, though not yet sufficient for a fully standardized large-scale assessment.
    Overall, the findings suggest that generative AI, when combined with ERI-based difficulty control and RAG-supported factual verification, can function as a supportive tool that enhances teachers’ capacity to develop literacy diagnostic assessments while still requiring professional human judgment at key stages. The study proposes a tentative R&D protocol for AI-assisted literacy test construction in school settings and provides empirical evidence of both its potential and its limitations. Given the constraints of a single school, a single grade level, and a small number of passages and items, further research is needed with more diverse samples, expanded item pools, and more sophisticated analytical approaches such as factor analysis, item response theory, and Rasch modeling.
    번역하기

    This study aimed to explore a development procedure for a classroom-based literacy diagnostic assessment for seventh-grade students by integrating generative AI, the EBS Reading Index (ERI), and Retrieval-Augmented Generation (RAG) technology. Drawing...

    This study aimed to explore a development procedure for a classroom-based literacy diagnostic assessment for seventh-grade students by integrating generative AI, the EBS Reading Index (ERI), and Retrieval-Augmented Generation (RAG) technology. Drawing on the framework of the EBS Literacy Diagnostic Assessment and the 2022 revised national curriculum, the study first specified passage design conditions (domain, topic, key concepts, text type, length) for five domains—arts, social studies, science, humanities, and integrated content. Based on these specifications, a generative AI model (ChatGPT) was used to produce draft passages and 15 multiple-choice items (three items per passage: literal comprehension, vocabulary, and inference/application). The AI-generated texts and items were then reviewed and revised through a two-step quality control process: factual and consistency checks using a RAG-based tool (NotebookLM) and subsequent human review by the researcher. To adjust text difficulty to the target grade level, ERI scores were calculated by combining quantitative indices derived from the National Institute of Korean Language’s vocabulary level lists and sentence complexity measures with qualitative ratings provided by five in-service Korean language teachers; passages were revised so that their ERI scores fell within the recommended range for seventh grade (7.0–8.5).
    A pilot test (n = 27) and a main test (n = 199) were conducted with first-year middle school students in Jeonju, South Korea. Classical test theory analyses were carried out, including item difficulty (p-values), item–total correlations, and Cronbach’s alpha for internal consistency, and the relationship between ERI scores and students’ perceived difficulty (6-point Likert scale) was examined. The final five passages showed ERI scores ranging from 7.0 to 8.1, indicating an appropriate difficulty level for the target grade. In both the pilot and main administrations, the rank order of ERI scores across passages exactly matched the rank order of mean perceived difficulty, suggesting that ERI can serve as a practical indicator for calibrating the difficulty of AI-generated texts. Cronbach’s alpha for the 15-item test was .79 in the pilot study and .64 (standardized α = .65) in the main study, indicating an acceptable level of internal consistency for an exploratory classroom- and school-level diagnostic tool, though not yet sufficient for a fully standardized large-scale assessment.
    Overall, the findings suggest that generative AI, when combined with ERI-based difficulty control and RAG-supported factual verification, can function as a supportive tool that enhances teachers’ capacity to develop literacy diagnostic assessments while still requiring professional human judgment at key stages. The study proposes a tentative R&D protocol for AI-assisted literacy test construction in school settings and provides empirical evidence of both its potential and its limitations. Given the constraints of a single school, a single grade level, and a small number of passages and items, further research is needed with more diverse samples, expanded item pools, and more sophisticated analytical approaches such as factor analysis, item response theory, and Rasch modeling.

    더보기

    목차 (Table of Contents)

    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구 문제 5
    • 3. 용어 정의 6
    • 가. 문해력 6
    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구 문제 5
    • 3. 용어 정의 6
    • 가. 문해력 6
    • 나. 생성형 AI 6
    • 다. ERI 지수 7
    • Ⅱ. 이론적 배경 8
    • 1. 생성형 AI 기반 독해 텍스트 및 문항 생성 연구 동향 8
    • 가. 교육용 대규모 평가에서의 AI 생성 지문 활용 8
    • 나. 독해 교육평가용 텍스트 및 문항 자동 생성 9
    • 다. 생성형 AI 기반 독해 문항의 실증적 검증 9
    • 라. 선행 연구의 시사점과 본 연구의 위치 10
    • 2. EBS 문해력 진단평가 12
    • 가. EBS 문해력 진단평가 개념 및 개발 배경 12
    • 나. 지문 유형 12
    • 다. 문항 수 및 구성 방식 12
    • 3. ERI(EBS Reading Index) 13
    • 4. 국어 어휘 등급화 연구 14
    • 5. 검색 증강 생성 기술(RAG 기술) 15
    • Ⅲ. 연구 방법 17
    • 1. 연구 참여자 17
    • 2. 연구 절차 18
    • 가. 지문 개발 19
    • 나. 난이도 검토(ERI 지수 이용) 20
    • 다. 평가 문항 개발 24
    • 라. 설문 도구 제작 25
    • 마. 예비검사(파일럿 테스트) 실시 25
    • 바. 본검사 실시 26
    • 사. 결과 분석 26
    • Ⅳ. 연구 결과 27
    • 1. 지문 개발 및 난이도 검토 결과 23
    • 가. 생성형 AI 기반 지문 개발 개요 27
    • 나. 지문 예시 27
    • 다. NotebookLM 및 인적 검토를 통한 지문 검토 결과 29
    • 라. ERI 지수 산출 및 난이도 조정 결과 31
    • 2. 평가 문항 개발 결과 24
    • 가. 문항 구성 개요 33
    • 나. NotebookLM을 활용한 사전 검토 결과 34
    • 다. 예비검사 실시 및 신뢰도 결과 분석 37
    • 라. 본검사용 최종 문항 세트 39
    • 3. 진단평가 결과 분석 41
    • 가. 시행 개요 41
    • 나. 본검사 도구의 기술통계 및 신뢰도 분석 결과 41
    • 다. 지문별 ERI 지수와 체감 난이도 간 일치도 분석 43
    • Ⅴ. 논의 및 결론 45
    • 1. 논의 45
    • 2. 결론 48
    • 3. 제언 50
    • 참고문헌 52
    • 부록 54
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼