RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    LLM 기반 문장 임베딩을 이용한 2022 개정교육과정 분석 : 전기전자 교과군을 중심으로 = An Analysis of the 2022 Revised National Curriculum Using LLM-based Sentence Embeddings: Focusing on the Electrical and Electronics Vocational Subject Group

    한글로보기

    https://www.riss.kr/link?id=T17428876

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구는 LLM 기반 문장 임베딩을 활용하여 2022 개정 전기전자 교과군 교육과정의 성취기준을 정량적으로 분석하고, 교과 간 의미 구조와 위계성을 탐색하고자 하였다. 이를 위해 KoSimCSE-RoBERTa 임베딩을 적용하여 전기전자 교과군 54개 과목의 성취기준 총 5,858문장을 벡터화 한 뒤, UMAP 차원 축소와 HDBSCAN 비지도 클러스터링을 수행하였다. 또한 기존 키워드 중심의 TF-IDF + K-Means, Word2Vec 평균 벡터 모델을 비교군으로 설정하여 LLM 기반 문장 임베딩의 구조적 타당성을 검증하였다.
    분석 결과, LLM 임베딩 기반 모델은 Baseline 모델에 비해 성취기준 간 의미적 관계를 보다 정교하게 반영하였다. TF-IDF와 Word2Vec 결과에서는 성취기준이 중앙에 밀집되어 교과 간 구분이 불명확하였으나, LLM 임베딩 기반 시각화에서는 교과별 군집이 뚜렷하게 분리되어 전기·전자 계열의 산업별, 기술별 특성이 명확히 나타났다. UMAP 결과에서는 이론(전공일반)과 실무(전력 계통, 반도체) 등 전문 분야가 분절되면서도 로봇, 3D 프린터 등 응용 실무 교과가 중간 영역에 위치해 교과 간 부분적인 근접성이 있음을 보여주었다. 덴드로그램 분석에서는 전공일반과 전공실무 교과가 상위 단계에서 분리되었고, 실무 교과 내부에서 전기 중심 교과와 전자 중심 교과로 하위 분화되는 위계 구조가 나타났다. 또한 클러스터드 히트맵 결과, 교과 간 평균 유사도는 0.88 이상으로 전반적으로 높은 연계성을 보였으나 특정 교과에서는 산업별 전문 용어와 절차 중심의 언어로 인해 상대적으로 낮은 유사도가 확인되었다.
    이러한 결과는 LLM 기반 문장 임베딩이 단순 키워드 빈도 분석을 넘어 문장의 의미적 맥락을 정량적으로 해석할 수 있음을 입증한다. 본 연구는 교육과정 문서를 데이터로 접근하여 교과 간 구조를 시각적·수치적으로 탐색한 시도이며, 단위학교의 교육과정 편성·재구성 및 융합형 수업 설계에 기초 자료로 활용될 수 있다. 다만 분석 대상이 2022 개정 교육과정의 17개 전문교과 교과군 중 전기전자 교과군만을 대상으로 한정하여 연구하였다는 한계가 있다.
    번역하기

    본 연구는 LLM 기반 문장 임베딩을 활용하여 2022 개정 전기전자 교과군 교육과정의 성취기준을 정량적으로 분석하고, 교과 간 의미 구조와 위계성을 탐색하고자 하였다. 이를 위해 KoSimCSE-RoBER...

    본 연구는 LLM 기반 문장 임베딩을 활용하여 2022 개정 전기전자 교과군 교육과정의 성취기준을 정량적으로 분석하고, 교과 간 의미 구조와 위계성을 탐색하고자 하였다. 이를 위해 KoSimCSE-RoBERTa 임베딩을 적용하여 전기전자 교과군 54개 과목의 성취기준 총 5,858문장을 벡터화 한 뒤, UMAP 차원 축소와 HDBSCAN 비지도 클러스터링을 수행하였다. 또한 기존 키워드 중심의 TF-IDF + K-Means, Word2Vec 평균 벡터 모델을 비교군으로 설정하여 LLM 기반 문장 임베딩의 구조적 타당성을 검증하였다.
    분석 결과, LLM 임베딩 기반 모델은 Baseline 모델에 비해 성취기준 간 의미적 관계를 보다 정교하게 반영하였다. TF-IDF와 Word2Vec 결과에서는 성취기준이 중앙에 밀집되어 교과 간 구분이 불명확하였으나, LLM 임베딩 기반 시각화에서는 교과별 군집이 뚜렷하게 분리되어 전기·전자 계열의 산업별, 기술별 특성이 명확히 나타났다. UMAP 결과에서는 이론(전공일반)과 실무(전력 계통, 반도체) 등 전문 분야가 분절되면서도 로봇, 3D 프린터 등 응용 실무 교과가 중간 영역에 위치해 교과 간 부분적인 근접성이 있음을 보여주었다. 덴드로그램 분석에서는 전공일반과 전공실무 교과가 상위 단계에서 분리되었고, 실무 교과 내부에서 전기 중심 교과와 전자 중심 교과로 하위 분화되는 위계 구조가 나타났다. 또한 클러스터드 히트맵 결과, 교과 간 평균 유사도는 0.88 이상으로 전반적으로 높은 연계성을 보였으나 특정 교과에서는 산업별 전문 용어와 절차 중심의 언어로 인해 상대적으로 낮은 유사도가 확인되었다.
    이러한 결과는 LLM 기반 문장 임베딩이 단순 키워드 빈도 분석을 넘어 문장의 의미적 맥락을 정량적으로 해석할 수 있음을 입증한다. 본 연구는 교육과정 문서를 데이터로 접근하여 교과 간 구조를 시각적·수치적으로 탐색한 시도이며, 단위학교의 교육과정 편성·재구성 및 융합형 수업 설계에 기초 자료로 활용될 수 있다. 다만 분석 대상이 2022 개정 교육과정의 17개 전문교과 교과군 중 전기전자 교과군만을 대상으로 한정하여 연구하였다는 한계가 있다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study aimed to quantitatively analyze the achievement standards of the Electrical and Electronics Vocational Subject Group in the 2022 Revised National Curriculum using LLM-based sentence embeddings and to explore the semantic structure and hierarchical relationships among subjects. To this end, 5,858 achievement-standard statements across 54 subjects were vectorized using KoSimCSE-RoBERTa, followed by UMAP dimensionality reduction and HDBSCAN unsupervised clustering. Additionally, traditional keyword-based models—TF-IDF with K-Means and Word2Vec average vectors—were adopted as baseline models to
    validate the structural validity of the LLM-based embeddings.
    The results showed that the LLM embedding–based model captured semantic relationships among achievement standards more precisely than the baseline models. While TF-IDF and Word2Vec produced dense, centrally clustered
    distributions in which subject boundaries were difficult to distinguish, the LLM-based visualizations revealed clearly separated clusters representing industrial and technological characteristics specific to the electrical and electronics domains.
    The UMAP visualization further showed that theoretical subjects (general subjects) and specialized practical areas (e.g., power systems, semiconductors) were distinctly segmented, whereas applied practice subjects such as robotics and 3D printing were positioned between clusters, suggesting partial inter-subject relatedness.
    Hierarchical clustering via dendrograms indicated that general subjects and practical subjects diverged at higher levels, with the practical subjects further branching into electrical-focused and electronics-focused subgroups. Clustered heatmap analysis showed overall high inter-subject similarity (average similarity ≥ 0.88), although certain subjects exhibited lower similarity due to the use of industry-specific terminology and procedure-oriented language.
    These findings demonstrate that LLM-based sentence embeddings can quantitatively capture the semantic context of curriculum statements beyond what is possible through simple keyword-frequency approaches. This study contributes to data-driven curriculum analysis by visualizing and numerically exploring inter-subject structures, offering foundational insights for curriculum organization, reconstruction, and the design of interdisciplinary or convergence-oriented instruction at the school level. However, the analysis is limited in scope, as it focuses solely on the Electrical and Electronics Vocational Subject Group among the 17 vocational subject groups included in the 2022 Revised National Curriculum.
    번역하기

    This study aimed to quantitatively analyze the achievement standards of the Electrical and Electronics Vocational Subject Group in the 2022 Revised National Curriculum using LLM-based sentence embeddings and to explore the semantic structure and hiera...

    This study aimed to quantitatively analyze the achievement standards of the Electrical and Electronics Vocational Subject Group in the 2022 Revised National Curriculum using LLM-based sentence embeddings and to explore the semantic structure and hierarchical relationships among subjects. To this end, 5,858 achievement-standard statements across 54 subjects were vectorized using KoSimCSE-RoBERTa, followed by UMAP dimensionality reduction and HDBSCAN unsupervised clustering. Additionally, traditional keyword-based models—TF-IDF with K-Means and Word2Vec average vectors—were adopted as baseline models to
    validate the structural validity of the LLM-based embeddings.
    The results showed that the LLM embedding–based model captured semantic relationships among achievement standards more precisely than the baseline models. While TF-IDF and Word2Vec produced dense, centrally clustered
    distributions in which subject boundaries were difficult to distinguish, the LLM-based visualizations revealed clearly separated clusters representing industrial and technological characteristics specific to the electrical and electronics domains.
    The UMAP visualization further showed that theoretical subjects (general subjects) and specialized practical areas (e.g., power systems, semiconductors) were distinctly segmented, whereas applied practice subjects such as robotics and 3D printing were positioned between clusters, suggesting partial inter-subject relatedness.
    Hierarchical clustering via dendrograms indicated that general subjects and practical subjects diverged at higher levels, with the practical subjects further branching into electrical-focused and electronics-focused subgroups. Clustered heatmap analysis showed overall high inter-subject similarity (average similarity ≥ 0.88), although certain subjects exhibited lower similarity due to the use of industry-specific terminology and procedure-oriented language.
    These findings demonstrate that LLM-based sentence embeddings can quantitatively capture the semantic context of curriculum statements beyond what is possible through simple keyword-frequency approaches. This study contributes to data-driven curriculum analysis by visualizing and numerically exploring inter-subject structures, offering foundational insights for curriculum organization, reconstruction, and the design of interdisciplinary or convergence-oriented instruction at the school level. However, the analysis is limited in scope, as it focuses solely on the Electrical and Electronics Vocational Subject Group among the 17 vocational subject groups included in the 2022 Revised National Curriculum.

    더보기

    목차 (Table of Contents)

    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구문제 5
    • Ⅱ. 이론적 배경 6
    • 1. 전기전자 교과군의 구조와 성격 6
    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구문제 5
    • Ⅱ. 이론적 배경 6
    • 1. 전기전자 교과군의 구조와 성격 6
    • 2. 연구 관련 주요 용어 정리 8
    • 3. 전통적 텍스트 임베딩 기법과 한계 9
    • 4. 거대언어모델(LLM) 기반 문장 임베딩과 비지도 학습 9
    • 가. 임베딩(Embedding)의 개념과 필요성 9
    • 나. BERT의 등장과 의의 10
    • 다. BERT의 한계와 개선 필요성 11
    • 라. RoBERTa의 개발과 주요 특징 12
    • 마. KoSimCSE-RoBERTa의 등장과 본 연구와의 연계성 13
    • 바. 비지도학습과 HDBSCAN 클러스터링 14
    • 사. UMAP 차원 축소와 시각화 15
    • 5. 클러스터링 품질 평가 지표 17
    • 6. 선행연구 고찰 및 본 연구의 위치 18
    • Ⅲ. 연구방법 21
    • 1. 연구 개요 21
    • 2. 연구 대상 및 자료 수집 21
    • 3. 자료 전처리 22
    • 4. 분석 도구 및 환경 23
    • 5. 분석 절차 23
    • 가. Baseline 분석 23
    • 나. LLM 기반 분석 24
    • 다. 비교 및 해석 24
    • 6. 분석 결과 정리 및 해석 기준 26
    • Ⅳ. 연구결과 27
    • 1. LLM분석 모델의 타당성 검증 27
    • 가. 정량적 성능 비교: 군집 구조의 안정성 평가 27
    • 나. 시각적 구조 비교: UMAP(성취기준 군집화), 덴드로그램 29
    • 2. LLM 임베딩 기반 분석 35
    • 가. UMAP을 통한 교과군 의미 지도 분석 35
    • 나. 덴드로그램을 통한 교과 간 위계 분석 38
    • 다. 클러스터드 히트맵(Clustered Heatmap) 교과 간 전체 유사도 분석 40
    • 라. 데이터 기반 융합 수업 설계 프레임워크 42
    • Ⅴ. 결론 44
    • 1. 결론 및 논의 44
    • 2. 제언 45
    • 참고문헌 47
    • Abstract 51
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼