RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Knowledge Localization for Efficient and Safe Language Models = 효율적이고 안전한 언어 모델을 위한 지식 영역 국소화

    한글로보기

    https://www.riss.kr/link?id=T17450538

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large language models (LLMs) have achieved outstanding performance across a wide range of natural language understanding and generation tasks. However, their internal mechanisms remain largely opaque, raising concerns about their trustworthiness. This dissertation presents a unified analytical framework—UniK (Unified Framework for Knowledge Localization)—that quantifies the contributions of each neuron to specific knowledge within LLMs. Building upon this foundation, the dissertation explores four major applications. First, Decomposition of Experts (DoE) introduces an unplug-and-play paradigm that dynamically deactivates task-irrelevant neurons, achieving inference speed-ups with minimal performance loss. Second, CRISPR mitigates biases in instruction-following models by pruning bias-related neurons, demonstrating that only a small subset of neurons encodes significant social or cognitive biases. Third, KLUE addresses privacy concerns by faithfully erasing sensitive knowledge through privacy-related neuron updates, supported by the FaithUn benchmark that considers knowledge interconnectedness. Finally, using the knowledge localization framework, we show that existing unlearning methods often hide rather than genuinely erase knowledge. To address this issue, we propose a regularization term that suppresses such hiding signals and promotes authentic forgetting. Collectively, these studies establish a clearer understanding of how task-specific, bias, and privacy knowledge is distributed within LLMs. These findings demonstrate that knowledge localization not only enhances interpretability but also enables proactive control over model behavior.
    번역하기

    Large language models (LLMs) have achieved outstanding performance across a wide range of natural language understanding and generation tasks. However, their internal mechanisms remain largely opaque, raising concerns about their trustworthiness. This...

    Large language models (LLMs) have achieved outstanding performance across a wide range of natural language understanding and generation tasks. However, their internal mechanisms remain largely opaque, raising concerns about their trustworthiness. This dissertation presents a unified analytical framework—UniK (Unified Framework for Knowledge Localization)—that quantifies the contributions of each neuron to specific knowledge within LLMs. Building upon this foundation, the dissertation explores four major applications. First, Decomposition of Experts (DoE) introduces an unplug-and-play paradigm that dynamically deactivates task-irrelevant neurons, achieving inference speed-ups with minimal performance loss. Second, CRISPR mitigates biases in instruction-following models by pruning bias-related neurons, demonstrating that only a small subset of neurons encodes significant social or cognitive biases. Third, KLUE addresses privacy concerns by faithfully erasing sensitive knowledge through privacy-related neuron updates, supported by the FaithUn benchmark that considers knowledge interconnectedness. Finally, using the knowledge localization framework, we show that existing unlearning methods often hide rather than genuinely erase knowledge. To address this issue, we propose a regularization term that suppresses such hiding signals and promotes authentic forgetting. Collectively, these studies establish a clearer understanding of how task-specific, bias, and privacy knowledge is distributed within LLMs. These findings demonstrate that knowledge localization not only enhances interpretability but also enables proactive control over model behavior.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    거대 언어 모델은 대규모 매개변수를 기반으로 다수의 텍스트 데이터로 학습한 결과, 다양한 자연어 이해 및 생성 과제에서 뛰어난 성능을 달성하였다. 그러나, 언어 모델 내부 동작 메커니즘은 여전히 대부분 밝혀지지 않았으며, 모델이 생성한 출력에 대한 신뢰성 우려가 지속적으로 제기되어 왔다. 본 논문에서는, 언어 모델 내 개별 뉴런이 특정 지식에 기여하는 정도를 식별하고 정량화하는 통합 분석 프레임워크인 UniK(Unified Framework for Knowledge Localization)을 제시한다. 본 논문은 이를 기반으로 네 가지 주요 응용 사례를 탐구한다. 첫째, 우리는 전문가 분해 방법론(DoE)을 활용하여 작업 별 뉴런을 식별하고 이를 동적으로 활성화하는 unplug-and-play 패러다임을 제안하여, 성능 저하를 최소화하면서 추론 속도를 향상시킨다. 둘째, 우리는 편향성 관련 뉴런을 식별하고 이를 가지치기함으로써 언어모델의 편향을 완화하는 CRISPR 방법론을 제안한다. 이를 기반으로 한 분석을 통해, 우리는 사회적・인지적 편향이 언어 모델 내 소수의 뉴런에 의해 주로 발현됨을 실험적으로 증명하였다. 셋째, 우리는 개인정보 관련 뉴런만을 선택적으로 학습하여 민감한 지식을 충실히 제거하는 KLUE 방법론을 제안한다. 마지막으로, 우리는 지식 제거 과정에 대한 심층 분석을 통해 기존 지식 제거 방법론들이 목표 지식을 제거하기보다는 은폐하려는 경향이 있음을 밝힌다. 또한, 지식을 숨기도록 유도하는 신호를 억제하는 정규화 학습 방법론을 제안하여 실질적인 망각을 달성하는 전략을 제안한다. 종합적으로, 본 연구는 다양한 종류의 지식 정보가 언어 모델 내에서 어떻게 분포하는 지를 포괄적으로 규명한다. 연구 결과, 우리는 지식 영역 국소화가 해석 가능성을 높일 뿐 아니라 모델의 행동을 능동적으로 제어하는 수단이 될 수 있음을 보여주었고 그 활용 가능성을 명확히 증명하였다.
    번역하기

    거대 언어 모델은 대규모 매개변수를 기반으로 다수의 텍스트 데이터로 학습한 결과, 다양한 자연어 이해 및 생성 과제에서 뛰어난 성능을 달성하였다. 그러나, 언어 모델 내부 동작 메커니...

    거대 언어 모델은 대규모 매개변수를 기반으로 다수의 텍스트 데이터로 학습한 결과, 다양한 자연어 이해 및 생성 과제에서 뛰어난 성능을 달성하였다. 그러나, 언어 모델 내부 동작 메커니즘은 여전히 대부분 밝혀지지 않았으며, 모델이 생성한 출력에 대한 신뢰성 우려가 지속적으로 제기되어 왔다. 본 논문에서는, 언어 모델 내 개별 뉴런이 특정 지식에 기여하는 정도를 식별하고 정량화하는 통합 분석 프레임워크인 UniK(Unified Framework for Knowledge Localization)을 제시한다. 본 논문은 이를 기반으로 네 가지 주요 응용 사례를 탐구한다. 첫째, 우리는 전문가 분해 방법론(DoE)을 활용하여 작업 별 뉴런을 식별하고 이를 동적으로 활성화하는 unplug-and-play 패러다임을 제안하여, 성능 저하를 최소화하면서 추론 속도를 향상시킨다. 둘째, 우리는 편향성 관련 뉴런을 식별하고 이를 가지치기함으로써 언어모델의 편향을 완화하는 CRISPR 방법론을 제안한다. 이를 기반으로 한 분석을 통해, 우리는 사회적・인지적 편향이 언어 모델 내 소수의 뉴런에 의해 주로 발현됨을 실험적으로 증명하였다. 셋째, 우리는 개인정보 관련 뉴런만을 선택적으로 학습하여 민감한 지식을 충실히 제거하는 KLUE 방법론을 제안한다. 마지막으로, 우리는 지식 제거 과정에 대한 심층 분석을 통해 기존 지식 제거 방법론들이 목표 지식을 제거하기보다는 은폐하려는 경향이 있음을 밝힌다. 또한, 지식을 숨기도록 유도하는 신호를 억제하는 정규화 학습 방법론을 제안하여 실질적인 망각을 달성하는 전략을 제안한다. 종합적으로, 본 연구는 다양한 종류의 지식 정보가 언어 모델 내에서 어떻게 분포하는 지를 포괄적으로 규명한다. 연구 결과, 우리는 지식 영역 국소화가 해석 가능성을 높일 뿐 아니라 모델의 행동을 능동적으로 제어하는 수단이 될 수 있음을 보여주었고 그 활용 가능성을 명확히 증명하였다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • List of Tables vi
    • List of Figures ix
    • Introduction 1
    • Abstract i
    • Contents ii
    • List of Tables vi
    • List of Figures ix
    • Introduction 1
    • Background 8
    • Unified Framework for Knowledge Localization 13
    • Knowledge Localization for Dynamic Compression 16
    • Knowledge Localization for Mitigating Biases 36
    • Knowledge Localization for Unlearning Privacy Risks 53
    • Analyzing Unlearning Dynamics using Knowledge Localization 76
    • Discussion and Conclusion 94
    • Abstract (In Korean) 115
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼