RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    대형 언어 모델의 지식 저장 메커니즘과 보안 취약성에 관한 연구 = A Study on Knowledge Storage Mechanisms and Security Vulnerabilities in Large Language Models

    한글로보기

    https://www.riss.kr/link?id=T17557196

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large language models store and utilize factual knowledge beyond simple text generation. Understanding where such knowledge is stored within a model is important not only for interpretability, but also for knowledge editing and security analysis. Knowledge editing is based on selectively identifying and modifying specific factual knowledge stored inside a model. Therefore, identifying the layers and modules in which knowledge is concentrated is essential for locating effective editing points and mitigating side effects that may arise during the editing process. In addition, such knowledge storage locations are also important from a security perspective, as they may be exploited for attacks such as malicious factual knowledge injection or backdoor insertion. However, prior studies have primarily focused on GPT-family architectures, and it remains insufficiently understood whether these knowledge storage mechanisms appear similarly across different autoregressive Transformer models. This study quantitatively analyzes the knowledge storage mechanisms of GPT-2-XL, LLaMA-3.2-1B, Qwen-2.5-1.5B, and DeepSeek-R1-Distill-Qwen-1.5B. To this end, we apply three causal analysis methods, Restoration Effect, Severing Effect, and Module Knockout, and compare the structural characteristics through which factual knowledge is stored across architectures. The results show that the storage mechanisms of factual knowledge differ across model architectures. GPT and LLaMA family models tend to store factual knowledge primarily in early MLP layers, whereas Qwen and DeepSeek family models rely more heavily on early Attention layers. In addition, Gini coefficient analysis reveals that these Attention-centered contributions are concentrated in specific layers. Furthermore, Module Knockout provides a more reliable validation of functional importance than Severing Effect alone, particularly in cases where such importance is difficult to confirm through Severing Effect. These findings extend our understanding of interpretability and the feasibility of knowledge editing in large language models and suggest the importance of considering architectural differences in the analysis of security vulnerabilities, including knowledge-editing-based attacks and backdoor insertion points.
    번역하기

    Large language models store and utilize factual knowledge beyond simple text generation. Understanding where such knowledge is stored within a model is important not only for interpretability, but also for knowledge editing and security analysis. Know...

    Large language models store and utilize factual knowledge beyond simple text generation. Understanding where such knowledge is stored within a model is important not only for interpretability, but also for knowledge editing and security analysis. Knowledge editing is based on selectively identifying and modifying specific factual knowledge stored inside a model. Therefore, identifying the layers and modules in which knowledge is concentrated is essential for locating effective editing points and mitigating side effects that may arise during the editing process. In addition, such knowledge storage locations are also important from a security perspective, as they may be exploited for attacks such as malicious factual knowledge injection or backdoor insertion. However, prior studies have primarily focused on GPT-family architectures, and it remains insufficiently understood whether these knowledge storage mechanisms appear similarly across different autoregressive Transformer models. This study quantitatively analyzes the knowledge storage mechanisms of GPT-2-XL, LLaMA-3.2-1B, Qwen-2.5-1.5B, and DeepSeek-R1-Distill-Qwen-1.5B. To this end, we apply three causal analysis methods, Restoration Effect, Severing Effect, and Module Knockout, and compare the structural characteristics through which factual knowledge is stored across architectures. The results show that the storage mechanisms of factual knowledge differ across model architectures. GPT and LLaMA family models tend to store factual knowledge primarily in early MLP layers, whereas Qwen and DeepSeek family models rely more heavily on early Attention layers. In addition, Gini coefficient analysis reveals that these Attention-centered contributions are concentrated in specific layers. Furthermore, Module Knockout provides a more reliable validation of functional importance than Severing Effect alone, particularly in cases where such importance is difficult to confirm through Severing Effect. These findings extend our understanding of interpretability and the feasibility of knowledge editing in large language models and suggest the importance of considering architectural differences in the analysis of security vulnerabilities, including knowledge-editing-based attacks and backdoor insertion points.

    더보기

    목차 (Table of Contents)

    • 제1장 서론 1
    • 제1절 연구 필요성 1
    • 제2절 연구 목적 및 연구 방법 2
    • 제3절 논문의 구성 3
    • 제2장 이론적 배경 5
    • 제1장 서론 1
    • 제1절 연구 필요성 1
    • 제2절 연구 목적 및 연구 방법 2
    • 제3절 논문의 구성 3
    • 제2장 이론적 배경 5
    • 제1절 Transformer 기반 언어 모델 5
    • 제2절 사실적 지식과 저장 메커니즘 8
    • 제3절 백도어 공격과 지식 저장 구조의 취약성 12
    • 제3장 제안 방법 15
    • 제1절 분석 방법 개요 15
    • 제2절 Restoration Effect 기반 분석 15
    • 제3절 Severing Effect 기반 분석 17
    • 제4절 Module Knockout 기반 분석 18
    • 제4장 실험 결과 및 분석 20
    • 제1절 실험 설정 20
    • 제2절 평가 지표 21
    • 제3절 Restoration Effect 결과 분석 25
    • 제4절 Severing Effect 결과 분석 30
    • 제5절 Module Knockout 결과 분석 35
    • 제5장 결론 38
    • 제6장 향후 연구 방향 40
    • 참고문헌 41
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼