RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Expert Subnetworks for Data- and Compute-Efficient Language Models = 언어 모델의 데이터 및 계산 효율성을 위한 전문가 서브네트워크 활용 연구

    한글로보기

    https://www.riss.kr/link?id=T17450301

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    언어 모델은 사회 전반에 막대한 영향을 끼치고 있지만, 그 혜택이 균등하게 분배되지는 않는다. 예를 들어, 저자원 언어는 학습 데이터가 부족할 뿐 아니라, 해당 언어를 사용하는 사회는 최신 모델을 사전학습, 적응, 서비스할 계산 자원이 부족한 경우가 많다.
    따라서 데이터와 계산 효율을 모두 갖춘 언어 모델이 요구된다.
    이 과제를 해결하기 위해 우리는 전문가 서브네트워크 (expert subnetwork)-- 모델 내부의 특화된 파라미터 하위집합-- 를 활용하여 데이터와 계산 효율을 동시에 향상시키는 방법을 제안한다.
    첫째, 다중언어 사전학습 단계에서 언어별 서브네트워크를 미리 지정하면 언어 간 부정적 간섭을 완화하여 각 언어 성능이 향상됨을 보인다.
    둘째, 계산 자원이 더욱 제한된 상황을 가정하여, 사전학습 대신 기존 모델을 새로운 언어로 적응하는 시나리오를 다룬다. 데이터 효율을 위해 원문 코퍼스와 음차 코퍼스를 혼합하고, 각 스크립트에 별도 서브네트워크를 할당한 뒤 표현을 융합하여 간섭을 줄인다.
    셋쨰, 추론 단계만 가능한 환경에서는 입력별로 적절한 서브네트워크를 동적으로 활성화하여 계산량을 최소화한 계산 효율형 LM을 구현한다.
    마지막으로, 데이터와 계산 효율을 동시에 극대화하기 위해 대규모 전문가 혼합 (MoE) 언어 모델에서 정적으로 서브네트워크를 선택하여 메모리 오버헤드를 줄인다.
    종합하면, 전문 서브네트워크를 활용해 데이터와 계산 측면 모두에서 효율적인 언어 모델을 구축할 수 있음을 입증하였다. 본 연구 결과는 전 세계적으로 포용적인 언어 기술을 실현하기 위한 실질적 기반을 제공한다.
    번역하기

    언어 모델은 사회 전반에 막대한 영향을 끼치고 있지만, 그 혜택이 균등하게 분배되지는 않는다. 예를 들어, 저자원 언어는 학습 데이터가 부족할 뿐 아니라, 해당 언어를 사용하는 사회는 ...

    언어 모델은 사회 전반에 막대한 영향을 끼치고 있지만, 그 혜택이 균등하게 분배되지는 않는다. 예를 들어, 저자원 언어는 학습 데이터가 부족할 뿐 아니라, 해당 언어를 사용하는 사회는 최신 모델을 사전학습, 적응, 서비스할 계산 자원이 부족한 경우가 많다.
    따라서 데이터와 계산 효율을 모두 갖춘 언어 모델이 요구된다.
    이 과제를 해결하기 위해 우리는 전문가 서브네트워크 (expert subnetwork)-- 모델 내부의 특화된 파라미터 하위집합-- 를 활용하여 데이터와 계산 효율을 동시에 향상시키는 방법을 제안한다.
    첫째, 다중언어 사전학습 단계에서 언어별 서브네트워크를 미리 지정하면 언어 간 부정적 간섭을 완화하여 각 언어 성능이 향상됨을 보인다.
    둘째, 계산 자원이 더욱 제한된 상황을 가정하여, 사전학습 대신 기존 모델을 새로운 언어로 적응하는 시나리오를 다룬다. 데이터 효율을 위해 원문 코퍼스와 음차 코퍼스를 혼합하고, 각 스크립트에 별도 서브네트워크를 할당한 뒤 표현을 융합하여 간섭을 줄인다.
    셋쨰, 추론 단계만 가능한 환경에서는 입력별로 적절한 서브네트워크를 동적으로 활성화하여 계산량을 최소화한 계산 효율형 LM을 구현한다.
    마지막으로, 데이터와 계산 효율을 동시에 극대화하기 위해 대규모 전문가 혼합 (MoE) 언어 모델에서 정적으로 서브네트워크를 선택하여 메모리 오버헤드를 줄인다.
    종합하면, 전문 서브네트워크를 활용해 데이터와 계산 측면 모두에서 효율적인 언어 모델을 구축할 수 있음을 입증하였다. 본 연구 결과는 전 세계적으로 포용적인 언어 기술을 실현하기 위한 실질적 기반을 제공한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Language models (LMs) are making huge impact to the society, yet their benefits remain unevenly distributed: low-resourced languages (LRLs) suffer from limited training data and societies that speak them often lack the compute to pre-train, adapt, or serve modern models.
    This calls language models to be more data- and compute-efficient, so that they can be utilized by societies with limited computational resources or data availability.
    To address this challenge, we propose leveraging expert subnetworks-- specialized subsets of parameters within LMs-- to enhance both data and compute efficiency.
    We explore expert subnetworks in three stages of LM lifecycle: pretraining, adaptation, and inference.
    We first show that pre-determining language-specific subnetworks in multilingual pre-training can mitigate negative interference between languages, leading to per-language performance boost.
    Second, we restrict the amount computation availability-- adapting an existing LM to a new expertise, such as languages unseen during pre-training, is often more feasible than pre-training from scratch. For data-efficiency, we mix the original and the transliterated corpora, allocate separate subnetworks to each script, and fuse their representations, to reduce the negative interference between the two.
    Third, we constrain the scenario to only inference-capable computation, and desire compute-efficient LM with expert subnetworks-- we dynamically activate subnetworks appropriate for given input.
    Finally, we push the limit of data- and compute-efficiency in the combined scenario: we statically select a subnetwork from a large mixture-of-experts (MoE) LM, to reduce the memory overhead of serving such model.
    Together, these results demonstrate that data- and compute- efficient LMs with expert subnetworks. Our findings lay practical groundwork for globally inclusive language technology.
    번역하기

    Language models (LMs) are making huge impact to the society, yet their benefits remain unevenly distributed: low-resourced languages (LRLs) suffer from limited training data and societies that speak them often lack the compute to pre-train, adapt, or ...

    Language models (LMs) are making huge impact to the society, yet their benefits remain unevenly distributed: low-resourced languages (LRLs) suffer from limited training data and societies that speak them often lack the compute to pre-train, adapt, or serve modern models.
    This calls language models to be more data- and compute-efficient, so that they can be utilized by societies with limited computational resources or data availability.
    To address this challenge, we propose leveraging expert subnetworks-- specialized subsets of parameters within LMs-- to enhance both data and compute efficiency.
    We explore expert subnetworks in three stages of LM lifecycle: pretraining, adaptation, and inference.
    We first show that pre-determining language-specific subnetworks in multilingual pre-training can mitigate negative interference between languages, leading to per-language performance boost.
    Second, we restrict the amount computation availability-- adapting an existing LM to a new expertise, such as languages unseen during pre-training, is often more feasible than pre-training from scratch. For data-efficiency, we mix the original and the transliterated corpora, allocate separate subnetworks to each script, and fuse their representations, to reduce the negative interference between the two.
    Third, we constrain the scenario to only inference-capable computation, and desire compute-efficient LM with expert subnetworks-- we dynamically activate subnetworks appropriate for given input.
    Finally, we push the limit of data- and compute-efficiency in the combined scenario: we statically select a subnetwork from a large mixture-of-experts (MoE) LM, to reduce the memory overhead of serving such model.
    Together, these results demonstrate that data- and compute- efficient LMs with expert subnetworks. Our findings lay practical groundwork for globally inclusive language technology.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Chapter 1 Introduction 1
    • Chapter 2 Background 7
    • Abstract i
    • Chapter 1 Introduction 1
    • Chapter 2 Background 7
    • 2.1 Expert Subnetworks for LMs 7
    • 2.2 Expert Subnetworks for Data- and Compute-Efficient LMs 8
    • 2.2.1 Pretraining: Negative Interference 9
    • 2.2.2 Adaptation: Data-Scarcity in LRLs 10
    • 2.2.3 Inference: Restricted Dynamic Subnetworks 11
    • 2.2.4 Combination: Large MoE Size 12
    • Chapter 3 Pretraining: Mitigating Interference in Multilingual Language Model Pretraining via Expert Subnetworks 15
    • 3.1 Multilingual Lottery Tickets with Zero-Shot NAS: Motivation and Overview 15
    • 3.1.1 Scale-then-Search Multilingual Tickets 17
    • 3.1.2 Zero-Shot NAS 19
    • 3.1.3 On Stably Measuring Negative Interference 20
    • 3.2 Experiments 22
    • 3.2.1 Experimental Settings 22
    • 3.2.2 Experimental Results 26
    • 3.2.3 Analysis: Ticket Similarity and Language Relatedness 28
    • Chapter 4 Adaptation: Per-Script Subnetworks for Language Model Adaptation to Unseen Low-Resource Language 31
    • 4.1 ScriptMix: Motivation and Overview 31
    • 4.1.1 Preliminaries: TL vs VA on MAD-X 33
    • 4.1.2 Dual-Script Corpus, the Gradient Conflict 35
    • 4.1.3 Language Module Separation 37
    • 4.1.4 AdapterFusion+: Reducing Inductive Bias 37
    • 4.1.5 Generalizing ScriptMix for UniPELT 38
    • 4.2 Experiments 42
    • 4.2.1 Experimental Settings 42
    • 4.2.2 Experimental Results and Analysis 46
    • 4.2.3 Analysis: Complementarity Visualization 49
    • Chapter 5 Inference: Efficient Inference Leveraging Dynamic Experts 50
    • 5.1 Generalized MoEfication: Motivation and Overview 50
    • 5.2 Preliminaries: MoEfication 53
    • 5.2.1 MoEfication under Hard Sparsity 53
    • 5.2.2 Utilizing Sparsity to Construct Experts 54
    • 5.2.3 Expert Selection 55
    • 5.3 Proposed Method 56
    • 5.3.1 Soft Sparsity for MoEfication in General Activation Functions 57
    • 5.3.2 Optimizing Sparsity with Representative Value 57
    • 5.3.3 Expert Selection 60
    • 5.4 Experiments 60
    • 5.4.1 Experimental Settings 60
    • 5.4.2 Experimental Results 62
    • Chapter 6 Pretrain+Adapt+Inference: Enabling MoE Serving with Expert Pruning 69
    • 6.1 Structured-Then-UNstructured Pruning (STUN): Motivation and Overview 69
    • 6.2 Preliminaries: MoE 72
    • 6.3 Expert-level Structured Pruning with O(1) GPU calls 73
    • 6.3.1 O(k^n/√n): Combinatorial Reconstruction Loss 74
    • 6.3.2 Towards O(n): Probabilistic Interpretation 75
    • 6.3.3 Towards O(1): Taylor Approximation and Selective Reconstruction 77
    • 6.4 Unstructured Pruning on Expert-pruned Model 78
    • 6.5 Experiments 79
    • 6.5.1 Experimental Settings 79
    • 6.5.2 Experimental Results 81
    • 6.5.3 Cost Analysis 85
    • Chapter 7 Conclusion 86
    • 7.1 Summary 86
    • 7.2 Limitations 86
    • 7.3 Future Directions 87
    • Appendix A Appendix 128
    • A.1 Appendix of ScriptMix 128
    • A.1.1 Results on Validation Split 128
    • A.1.2 Standard Deviations of Experiments 128
    • A.1.3 Visualizing the Complementarity 128
    • A.2 Appendix of STUN 133
    • A.2.1 Derivation From O(k^n/√n) to O(1) 133
    • A.2.2 Implementation Details 137
    • A.2.3 Other Clustering Algorithms 138
    • A.2.4 Ablation Studies 138
    • A.2.5 Detailed Results for RQ2 140
    • A.2.6 Why GSM8K is Harder 140
    • Acknowledgements 142
    • 요약 145
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼