RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    내용 체계 범주에 따른 서술형 문항에 대한 자연어 생성 모델과 자연어 이해 모델의 자동 채점 결과 비교 연구 : 중학교 사회 경제 단원을 중심으로 = A Comparative Study on Automated Scoring of Constructed-Response Items Using Natural Language Generation and Understanding Models Based on Content System: Focusing on the Economics Unit in Middle School Social Studies

    한글로보기

    https://www.riss.kr/link?id=T17451335

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Descriptive assessments are an important evaluation method capable of assessing students' diverse competencies, such as creativity and higher-order thinking skills. However, their implementation in schools has faced significant challenges due to the substantial time and cost required for grading, coupled with difficulties in ensuring inter-rater reliability. Recently emerged artificial intelligence technology has shown new potential to overcome these limitations. Therefore, this study analyzed the effectiveness of an automated grading program utilizing the latest AI technologies—the natural language generation model (GPT model) and the natural language understanding model (BERT model)—as a method to reduce teachers' burden in expanding essay-type assessments in schools and to enhance assessment reliability. This could contribute to increasing the practical applicability of automated grading.
    Specifically, descriptive questions on the concept of ‘demand’ from the middle school social studies and economics unit were developed by categorizing them according to the content framework categories of the 2022 revised curriculum: ‘Knowledge and Understanding’, ‘Process and Function’, and ‘Values and Attitudes’. The automatic grading performance and characteristics of the natural language generation model and natural language understanding model were then compared for each question type. For this study, a total of 900 answer data points collected from 300 third-year middle school students at three middle schools in Seoul were utilized. Student answers were graded using an automatic grading program, and the results were compared and analyzed against teacher grading. The natural language generation model employed an automatic grading program based on Google Sheets, while the natural language understanding model utilized an automatic grading program built on Python within Google Colab for the research.
    The main findings of this study are as follows. First, we derived the optimal prompt type and automatic scoring method for each AI model. Natural language generation models showed the highest performance with the ‘few-shot’ type, which provides rubrics, example answers, and specific scoring cases. Natural language understanding models demonstrated high performance with the ‘classification-based’ method, which predicts score intervals by learning patterns in correct answer data. Second, the most effective AI model differed depending on the content domain category of the item. For items in the ‘Knowledge/Understanding’ content framework category, where correct answers are clear, both models showed a very high level of agreement with teacher grading (QWK ≥ 0.9). For items in the ‘Process/Function’ content framework category, requiring logical thinking, the natural language understanding model proved effective. For items in the ‘Value/Attitude’ content framework category, involving the judgment of diverse values in the affective domain, the natural language generation model was the effective model. Third, differences in grading tendencies emerged between models. The natural language generation model exhibited a strict tendency toward ‘undergrading’ due to the influence of hallucination prevention training. Conversely, the natural language understanding model showed a relatively lenient tendency toward ‘overgrading,’ reacting sensitively to whether keywords were present in student responses. These findings suggest that AI-powered automated grading can be an effective tool to assist teachers in grading within school settings. However, rather than applying a single AI model indiscriminately, it is necessary to selectively apply appropriate models and optimal strategies by considering the nature of the item according to the content framework category and the purpose of the assessment (e.g., strict selection or generous feedback).
    By proposing optimal AI automatic grading approaches for each content system category, this study is expected to contribute to enhancing the quality of essay-type assessments by reducing the burden teachers feel in schools. It also holds significance in providing foundational data for building AI-based assessment systems.
    However, the GPT and BERT models used in this study are based on the most recent models available at this time. The possibility that more effective strategies and methods may emerge in new assessment contexts when newer models are released in the future suggests the need for ongoing research in automated grading. Furthermore, since some scoring models require additional effort for teachers to use directly in school settings, further support will be needed to devise ways to utilize them in more convenient and user-friendly ways.
    번역하기

    Descriptive assessments are an important evaluation method capable of assessing students' diverse competencies, such as creativity and higher-order thinking skills. However, their implementation in schools has faced significant challenges due to the s...

    Descriptive assessments are an important evaluation method capable of assessing students' diverse competencies, such as creativity and higher-order thinking skills. However, their implementation in schools has faced significant challenges due to the substantial time and cost required for grading, coupled with difficulties in ensuring inter-rater reliability. Recently emerged artificial intelligence technology has shown new potential to overcome these limitations. Therefore, this study analyzed the effectiveness of an automated grading program utilizing the latest AI technologies—the natural language generation model (GPT model) and the natural language understanding model (BERT model)—as a method to reduce teachers' burden in expanding essay-type assessments in schools and to enhance assessment reliability. This could contribute to increasing the practical applicability of automated grading.
    Specifically, descriptive questions on the concept of ‘demand’ from the middle school social studies and economics unit were developed by categorizing them according to the content framework categories of the 2022 revised curriculum: ‘Knowledge and Understanding’, ‘Process and Function’, and ‘Values and Attitudes’. The automatic grading performance and characteristics of the natural language generation model and natural language understanding model were then compared for each question type. For this study, a total of 900 answer data points collected from 300 third-year middle school students at three middle schools in Seoul were utilized. Student answers were graded using an automatic grading program, and the results were compared and analyzed against teacher grading. The natural language generation model employed an automatic grading program based on Google Sheets, while the natural language understanding model utilized an automatic grading program built on Python within Google Colab for the research.
    The main findings of this study are as follows. First, we derived the optimal prompt type and automatic scoring method for each AI model. Natural language generation models showed the highest performance with the ‘few-shot’ type, which provides rubrics, example answers, and specific scoring cases. Natural language understanding models demonstrated high performance with the ‘classification-based’ method, which predicts score intervals by learning patterns in correct answer data. Second, the most effective AI model differed depending on the content domain category of the item. For items in the ‘Knowledge/Understanding’ content framework category, where correct answers are clear, both models showed a very high level of agreement with teacher grading (QWK ≥ 0.9). For items in the ‘Process/Function’ content framework category, requiring logical thinking, the natural language understanding model proved effective. For items in the ‘Value/Attitude’ content framework category, involving the judgment of diverse values in the affective domain, the natural language generation model was the effective model. Third, differences in grading tendencies emerged between models. The natural language generation model exhibited a strict tendency toward ‘undergrading’ due to the influence of hallucination prevention training. Conversely, the natural language understanding model showed a relatively lenient tendency toward ‘overgrading,’ reacting sensitively to whether keywords were present in student responses. These findings suggest that AI-powered automated grading can be an effective tool to assist teachers in grading within school settings. However, rather than applying a single AI model indiscriminately, it is necessary to selectively apply appropriate models and optimal strategies by considering the nature of the item according to the content framework category and the purpose of the assessment (e.g., strict selection or generous feedback).
    By proposing optimal AI automatic grading approaches for each content system category, this study is expected to contribute to enhancing the quality of essay-type assessments by reducing the burden teachers feel in schools. It also holds significance in providing foundational data for building AI-based assessment systems.
    However, the GPT and BERT models used in this study are based on the most recent models available at this time. The possibility that more effective strategies and methods may emerge in new assessment contexts when newer models are released in the future suggests the need for ongoing research in automated grading. Furthermore, since some scoring models require additional effort for teachers to use directly in school settings, further support will be needed to devise ways to utilize them in more convenient and user-friendly ways.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    서술형 평가는 학생의 창의력, 고차사고력 등의 다양한 역량을 평가할 수 있는 중요한 평가 방법이지만, 채점에 많은 시간과 비용이 들고 채점자 간 신뢰도 확보가 어렵다는 문제가 발생하기 때문에 학교 현장에서 활용하는 데 많은 어려움이 있었다. 최근 등장한 인공지능 기술은 이러한 한계를 극복할 수 있는 새로운 가능성을 보여주었다. 따라서 본 연구는 학교 현장의 서술형 평가 확대에 대한 교사의 부담을 줄여주고, 평가의 신뢰도를 높이기 위한 방법으로 최신 인공지능 기술인 자연어 생성 모델(GPT 모델)과 자연어 이해 모델(BERT 모델)을 활용한 자동 채점 프로그램의 효과성을 분석하였다. 이는 자동 채점의 실질적인 활용 가능성을 높이는 데 이바지할 수 있을 것이다.
    특히 중학교 사회 경제 단원의 ‘수요’ 개념에 대한 서술형 문항을 2022 개정 교육과정의 내용 체계 범주인 ‘지식·이해’, ‘과정·기능’, ‘가치·태도’로 구분하여 개발하고, 각 문항 유형에 따른 자연어 생성 모델과 자연어 이해 모델의 자동 채점 성능과 특징을 비교하였다. 연구를 위해 서울 소재 3개 중학교 3학년 학생 300명을 대상으로 수집된 총 900개의 답안 데이터를 활용하였다. 학생 답안을 자동 채점 프로그램으로 채점하였고, 이를 교사 채점 결과와 비교·분석하였다. 자연어 생성 모델은 구글 스프레드시트 기반의 자동 채점 프로그램을 활용하였고, 자연어 이해 모델은 구글 코랩에서 파이썬 기반의 자동 채점 프로그램을 구축하여 연구를 진행하였다.
    본 연구의 주요 결과는 다음과 같다. 첫째, 인공지능 모델별 최적 프롬프트 유형 및 자동 채점 방식을 도출하였다. 자연어 생성 모델은 루브릭과 예시 답안, 구체적인 채점 사례를 제공하는 ‘퓨샷’ 유형에서 가장 높은 성능을 보였으며, 자연어 이해 모델은 정답 데이터의 패턴을 학습하여 점수 급간을 예측하는 ‘분류 기반’ 방식이 높은 성능을 보였다. 둘째, 문항의 내용 체계 범주에 따라 효과적인 인공지능 모델이 다르게 나타났다. 정답이 명확한 ‘지식·이해’ 내용 체계 범주의 문항에서는 두 모델 모두 교사 채점과의 일치도가 매우 높은 수준(QWK 0.9 이상)으로 나타났다. 논리적인 사고가 필요한 ‘과정·기능’ 내용 체계 범주의 문항에서는 자연어 이해 모델이, 정의적 영역으로 다양한 가치를 판단하는 ‘가치·태도’ 내용 체계 범주의 문항에서는 자연어 생성 모델이 효과적인 모델로 나타났다. 셋째, 모델별 채점 성향의 차이가 나타났다. 자연어 생성 모델은 환각 방지 학습의 영향으로 엄격한 ‘과소 채점’ 경향을 보였지만, 자연어 이해 모델은 키워드가 학생 답안에 포함되었는지 여부에 민감하게 반응하여 상대적으로 관대한 ‘과대 채점’ 경향을 보였다. 이러한 연구 결과는 인공지능을 활용한 자동 채점이 학교 현장에 교사의 채점을 보조할 수 있는 효과적인 도구가 될 수 있지만, 하나의 인공지능 모델을 그대로 적용하기보다는 내용 체계 범주에 따른 문항의 성격과 평가의 목적(엄격한 선발 혹은 관대한 피드백 등)을 고려하여 적절한 모델과 최적의 전략을 선별적으로 적용해야 함을 시사한다.
    본 연구는 내용 체계 범주별로 최적의 인공지능 자동 채점 방안을 제시함으로써, 앞으로 인공지능을 보조 채점자로 활용한다면 학교 현장에서 교사가 느끼는 서술형 평가에 대한 부담을 줄여 서술형 평가 내실화에 기여할 수 있을 것으로 기대된다. 또한 인공지능 기반 평가 시스템 구축을 위한 기초 자료를 제공한다는 점에서 의의가 있다고 할 수 있다.
    다만, 본 연구가 사용한 GPT 모델과 BERT 모델은 현재 시점에서 가장 최신인 모델에 기반한 것으로, 향후 새로운 모델이 출시되면 새로운 평가 맥락에서 더 효과적인 전략과 방법이 존재할 수 있다는 점은 앞으로도 지속적인 자동 채점 연구가 필요하다는 것을 시사한다. 또한 일부 채점 모델은 교사가 학교 현장에서 그대로 사용하기에는 추가적인 노력이 필요하다는 점에서 좀 더 편리하고 쉬운 방식으로 활용할 수 있는 방안을 고안하기 위한 추가적인 지원이 필요할 것이다.
    번역하기

    서술형 평가는 학생의 창의력, 고차사고력 등의 다양한 역량을 평가할 수 있는 중요한 평가 방법이지만, 채점에 많은 시간과 비용이 들고 채점자 간 신뢰도 확보가 어렵다는 문제가 발생하...

    서술형 평가는 학생의 창의력, 고차사고력 등의 다양한 역량을 평가할 수 있는 중요한 평가 방법이지만, 채점에 많은 시간과 비용이 들고 채점자 간 신뢰도 확보가 어렵다는 문제가 발생하기 때문에 학교 현장에서 활용하는 데 많은 어려움이 있었다. 최근 등장한 인공지능 기술은 이러한 한계를 극복할 수 있는 새로운 가능성을 보여주었다. 따라서 본 연구는 학교 현장의 서술형 평가 확대에 대한 교사의 부담을 줄여주고, 평가의 신뢰도를 높이기 위한 방법으로 최신 인공지능 기술인 자연어 생성 모델(GPT 모델)과 자연어 이해 모델(BERT 모델)을 활용한 자동 채점 프로그램의 효과성을 분석하였다. 이는 자동 채점의 실질적인 활용 가능성을 높이는 데 이바지할 수 있을 것이다.
    특히 중학교 사회 경제 단원의 ‘수요’ 개념에 대한 서술형 문항을 2022 개정 교육과정의 내용 체계 범주인 ‘지식·이해’, ‘과정·기능’, ‘가치·태도’로 구분하여 개발하고, 각 문항 유형에 따른 자연어 생성 모델과 자연어 이해 모델의 자동 채점 성능과 특징을 비교하였다. 연구를 위해 서울 소재 3개 중학교 3학년 학생 300명을 대상으로 수집된 총 900개의 답안 데이터를 활용하였다. 학생 답안을 자동 채점 프로그램으로 채점하였고, 이를 교사 채점 결과와 비교·분석하였다. 자연어 생성 모델은 구글 스프레드시트 기반의 자동 채점 프로그램을 활용하였고, 자연어 이해 모델은 구글 코랩에서 파이썬 기반의 자동 채점 프로그램을 구축하여 연구를 진행하였다.
    본 연구의 주요 결과는 다음과 같다. 첫째, 인공지능 모델별 최적 프롬프트 유형 및 자동 채점 방식을 도출하였다. 자연어 생성 모델은 루브릭과 예시 답안, 구체적인 채점 사례를 제공하는 ‘퓨샷’ 유형에서 가장 높은 성능을 보였으며, 자연어 이해 모델은 정답 데이터의 패턴을 학습하여 점수 급간을 예측하는 ‘분류 기반’ 방식이 높은 성능을 보였다. 둘째, 문항의 내용 체계 범주에 따라 효과적인 인공지능 모델이 다르게 나타났다. 정답이 명확한 ‘지식·이해’ 내용 체계 범주의 문항에서는 두 모델 모두 교사 채점과의 일치도가 매우 높은 수준(QWK 0.9 이상)으로 나타났다. 논리적인 사고가 필요한 ‘과정·기능’ 내용 체계 범주의 문항에서는 자연어 이해 모델이, 정의적 영역으로 다양한 가치를 판단하는 ‘가치·태도’ 내용 체계 범주의 문항에서는 자연어 생성 모델이 효과적인 모델로 나타났다. 셋째, 모델별 채점 성향의 차이가 나타났다. 자연어 생성 모델은 환각 방지 학습의 영향으로 엄격한 ‘과소 채점’ 경향을 보였지만, 자연어 이해 모델은 키워드가 학생 답안에 포함되었는지 여부에 민감하게 반응하여 상대적으로 관대한 ‘과대 채점’ 경향을 보였다. 이러한 연구 결과는 인공지능을 활용한 자동 채점이 학교 현장에 교사의 채점을 보조할 수 있는 효과적인 도구가 될 수 있지만, 하나의 인공지능 모델을 그대로 적용하기보다는 내용 체계 범주에 따른 문항의 성격과 평가의 목적(엄격한 선발 혹은 관대한 피드백 등)을 고려하여 적절한 모델과 최적의 전략을 선별적으로 적용해야 함을 시사한다.
    본 연구는 내용 체계 범주별로 최적의 인공지능 자동 채점 방안을 제시함으로써, 앞으로 인공지능을 보조 채점자로 활용한다면 학교 현장에서 교사가 느끼는 서술형 평가에 대한 부담을 줄여 서술형 평가 내실화에 기여할 수 있을 것으로 기대된다. 또한 인공지능 기반 평가 시스템 구축을 위한 기초 자료를 제공한다는 점에서 의의가 있다고 할 수 있다.
    다만, 본 연구가 사용한 GPT 모델과 BERT 모델은 현재 시점에서 가장 최신인 모델에 기반한 것으로, 향후 새로운 모델이 출시되면 새로운 평가 맥락에서 더 효과적인 전략과 방법이 존재할 수 있다는 점은 앞으로도 지속적인 자동 채점 연구가 필요하다는 것을 시사한다. 또한 일부 채점 모델은 교사가 학교 현장에서 그대로 사용하기에는 추가적인 노력이 필요하다는 점에서 좀 더 편리하고 쉬운 방식으로 활용할 수 있는 방안을 고안하기 위한 추가적인 지원이 필요할 것이다.

    더보기

    목차 (Table of Contents)

    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구 문제 4
    • Ⅱ. 이론적 배경 5
    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구 문제 4
    • Ⅱ. 이론적 배경 5
    • 1. 2022 개정 사회과 교육과정과 내용 체계 5
    • 1) 2022 개정 교육과정 내용 체계 및 성취기준 특징 5
    • 2) 2022 개정 사회과 교육과정의 특징 6
    • 3) 사회과 교육과정 내용 체계 범주 7
    • 2. 서술형 평가 11
    • 1) 서술형 평가의 의의 11
    • 2) 일반사회 경제 영역에서의 서술형 평가 12
    • 3) 서술형 평가에 대한 교사의 인식 14
    • 4) 서술형 평가의 한계 및 자동 채점 연구 시행 배경 15
    • 3. 인공지능 기반 자동 채점 17
    • 1) 인공지능을 활용한 자동 채점 단계 17
    • 2) 트랜스포머 18
    • 3) 자연어 생성 모델(GPT 모델) 20
    • 4) 자연어 이해 모델(BERT 모델) 22
    • 5) GPT 모델과 BERT 모델의 비교 23
    • 4. 선행 연구 24
    • Ⅲ. 연구 방법 30
    • 1. 연구 대상 30
    • 2. 연구 교과 및 주제 30
    • 3. 연구 절차 31
    • 1) 자동 채점 프로그램 개발 37
    • 2) 데이터셋 구성 및 동질성 검증 47
    • 2) 자동 채점 프로그램 결과 비교 지표 48
    • Ⅳ. 연구 결과 54
    • 1. 자연어 생성 모델 프롬프트 엔지니어링 결과 54
    • 1) 서술형 문항 1(지식·이해) 54
    • 2) 서술형 문항 2(과정·기능) 56
    • 3) 서술형 문항 3(가치·태도) 58
    • 4) 소결 60
    • 2. 자연어 이해 모델 기반 자동 채점 방식 분석 결과 61
    • 1) 서술형 문항 1(지식·이해) 61
    • 2) 서술형 문항 2(과정·기능) 63
    • 3) 서술형 문항 3(가치·태도) 65
    • 4) 소결 67
    • 3. 자연어 생성 모델과 자연어 이해 모델 비교 분석 68
    • 1) 서술형 문항 1(지식·이해) 68
    • 2) 서술형 문항 2(과정·기능) 73
    • 3) 서술형 문항 3(가치·태도) 78
    • Ⅴ. 결론 84
    • 1. 연구 요약 및 결론 84
    • 2. 연구의 의의 및 제언 87
    • 참고문헌 91
    • 부록 100
    • Abstract 132
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼