RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Task-aware Block Pruning with Output Distribution Signals for Large Language Models = 대형 언어 모델을 위한 출력 분포 신호 기반 작업 지향형 트랜스포머 블록 가지치기

    한글로보기

    https://www.riss.kr/link?id=T17451074

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large language models (LLMs) provide excellent performance across a wide range of natural language tasks, but their practical deployment in real-world systems is often limited by significant inference costs and latency under strict resource constraints. While block pruning has emerged as an effective strategy to reduce computation and preserve the structural coherence of transformer architectures, existing methods typically rely on representation similarity or costly sensitivity analyses, which only weakly reflect task-aware model behavior and are difficult to scale in practice. In this thesis, Task-aware Block Pruning (TaBP) is proposed as a novel block pruning framework that quantifies block-level uncertainty by attaching lightweight LM heads to each block and measuring entropy-based statistics of their early-exited output distributions on a task-specific calibration dataset, in order to identify prunable blocks. Extensive experiments on multiple LLM backbones and downstream benchmarks validate the effectiveness of the proposed method, demonstrating substantial efficiency gains without compromising task performance, while avoiding computationally expensive sensitivity analyses.
    번역하기

    Large language models (LLMs) provide excellent performance across a wide range of natural language tasks, but their practical deployment in real-world systems is often limited by significant inference costs and latency under strict resource constraint...

    Large language models (LLMs) provide excellent performance across a wide range of natural language tasks, but their practical deployment in real-world systems is often limited by significant inference costs and latency under strict resource constraints. While block pruning has emerged as an effective strategy to reduce computation and preserve the structural coherence of transformer architectures, existing methods typically rely on representation similarity or costly sensitivity analyses, which only weakly reflect task-aware model behavior and are difficult to scale in practice. In this thesis, Task-aware Block Pruning (TaBP) is proposed as a novel block pruning framework that quantifies block-level uncertainty by attaching lightweight LM heads to each block and measuring entropy-based statistics of their early-exited output distributions on a task-specific calibration dataset, in order to identify prunable blocks. Extensive experiments on multiple LLM backbones and downstream benchmarks validate the effectiveness of the proposed method, demonstrating substantial efficiency gains without compromising task performance, while avoiding computationally expensive sensitivity analyses.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    대형 언어 모델은 다양한 자연어 처리 과제에서 우수한 성능을 보이지만, 막대한 추론 비용 때문에 실제 서비스 적용에는 여전히 제약이 따른다. 블록 프루닝(block pruning)은 모델의 구조적 일관성을 유지하면서 지연 시간을 줄이는 효과적인 방법이지만, 기존 기법들은 주로 표현 벡터의 유사도나 고비용 민감도 분석에 의존하여 작업 특이적인 모델 행동을 충분히 반영하지 못한다. 본 논문에서는 출력 분포의 엔트로피 기반 추정치를 활용해 중요도가 낮은 트랜스포머 블록을 정밀하게 식별하는 출력 지향형 블록 프루닝 방법을 제안한다. 다양한 벤치마크에 대한 실험 결과, 제안 기법은 다운스트림 작업 성능을 유지하면서도 추론 효율을 크게 향상시킬 수 있음을 보인다.
    번역하기

    대형 언어 모델은 다양한 자연어 처리 과제에서 우수한 성능을 보이지만, 막대한 추론 비용 때문에 실제 서비스 적용에는 여전히 제약이 따른다. 블록 프루닝(block pruning)은 모델의 구조적 일...

    대형 언어 모델은 다양한 자연어 처리 과제에서 우수한 성능을 보이지만, 막대한 추론 비용 때문에 실제 서비스 적용에는 여전히 제약이 따른다. 블록 프루닝(block pruning)은 모델의 구조적 일관성을 유지하면서 지연 시간을 줄이는 효과적인 방법이지만, 기존 기법들은 주로 표현 벡터의 유사도나 고비용 민감도 분석에 의존하여 작업 특이적인 모델 행동을 충분히 반영하지 못한다. 본 논문에서는 출력 분포의 엔트로피 기반 추정치를 활용해 중요도가 낮은 트랜스포머 블록을 정밀하게 식별하는 출력 지향형 블록 프루닝 방법을 제안한다. 다양한 벤치마크에 대한 실험 결과, 제안 기법은 다운스트림 작업 성능을 유지하면서도 추론 효율을 크게 향상시킬 수 있음을 보인다.

    더보기

    목차 (Table of Contents)

    • Chapter 1. Introduction 4
    • Chapter 2. Literature Review 7
    • Chapter 3. Methodology 15
    • Chapter 4. Experiments 22
    • Chapter 5. Results 24
    • Chapter 1. Introduction 4
    • Chapter 2. Literature Review 7
    • Chapter 3. Methodology 15
    • Chapter 4. Experiments 22
    • Chapter 5. Results 24
    • Chapter 6. Conclusion 32
    • Bibliography 33
    • Abstract in Korean 47
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼