RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Analysis of NoPE Replacement for Rotational Positional Embeddings in Transformer Architectures = 트랜스포머 위치 임베딩 제거의 영향 분석

    한글로보기

    https://www.riss.kr/link?id=T17449798

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Rotary positional embedding (RoPE) is widely used in modern large language models, but recent evidence
    suggests it can introduce inductive biases that hinder long-context reasoning, such as attention sink and
    degraded mid-context sensitivity. In this work, we examine whether RoPE can be partially removed in
    pretrained models—by converting selected attention heads to No Positional Encoding (NoPE)—without any
    additional training.
    Through systematic head- and layer-wise ablations on Llama-3.2-3B-Instruct evaluated with the RULER
    benchmark at 4K–8K context lengths, we find that NoPE replacement is not uniformly harmful. While
    early layers are highly sensitive to RoPE removal, later layers contain heads for which NoPE replacement
    consistently improves performance. Both position–awareness–guided and random head selection can yield
    gains, indicating that benefits arise from reducing positional bias rather than precise head identification.
    Finally, attention analysis shows that NoPE replacement mitigates attention sink by redistributing attention
    toward mid-context tokens. These results demonstrate that selective NoPE replacement at inference time can
    improve long-context behavior in pretrained models without retraining.
    번역하기

    Rotary positional embedding (RoPE) is widely used in modern large language models, but recent evidence suggests it can introduce inductive biases that hinder long-context reasoning, such as attention sink and degraded mid-context sensitivity. In this ...

    Rotary positional embedding (RoPE) is widely used in modern large language models, but recent evidence
    suggests it can introduce inductive biases that hinder long-context reasoning, such as attention sink and
    degraded mid-context sensitivity. In this work, we examine whether RoPE can be partially removed in
    pretrained models—by converting selected attention heads to No Positional Encoding (NoPE)—without any
    additional training.
    Through systematic head- and layer-wise ablations on Llama-3.2-3B-Instruct evaluated with the RULER
    benchmark at 4K–8K context lengths, we find that NoPE replacement is not uniformly harmful. While
    early layers are highly sensitive to RoPE removal, later layers contain heads for which NoPE replacement
    consistently improves performance. Both position–awareness–guided and random head selection can yield
    gains, indicating that benefits arise from reducing positional bias rather than precise head identification.
    Finally, attention analysis shows that NoPE replacement mitigates attention sink by redistributing attention
    toward mid-context tokens. These results demonstrate that selective NoPE replacement at inference time can
    improve long-context behavior in pretrained models without retraining.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    Rotary positional embedding (RoPE)은현대의대규모 언어 모델에서널리 사용되고있으나, 최근연구
    들은RoPE가attention sink나 중간문맥 감도저하와같은장문 문맥 추론을방해하는 inductive bias를
    유발할 수있음을시사한다. 본연구에서는 추가적인학습없이, 사전학습된모델내 일부 attention head
    를 No Positional Encoding (NoPE)으로변환함으로써 RoPE를 부분적으로제거할 수있는지를 분석한다.
    Llama-3.2-3B-Instruct 모델을대상으로RULER 벤치마크에서4K–8K 문맥 길이범위에 걸쳐 head- 및
    layer-단위의체계적인ablation 실험을수행한 결과, NoPE 변환이항상 성능 저하를 초래하지는 않음을
    확인하였다. 초기 레이어는 RoPE 제거에 매우민감한 반면, 후반레이어에는 NoPE로대체했을때 일관된
    성능 향상을보이는 attention head들이존재한다. 또한 position-awareness-guided 선택과무작위head 선택
    모두에서성능 향상이관찰되어, 이러한 이점이정확한 head 식별보다는 positional bias를 완화하는 데서
    비롯됨을시사한다.
    마지막으로attention 분석을통해 NoPE 변환이attention을중간문맥 토큰으로재분배함으로써 attention
    sink 현상을완화함을보인다. 이러한 결과는 재학습없이도inference 단계에서선택적으로NoPE를 적용
    하는 것이사전학습된모델의장문 문맥 처리 능력을향상시킬 수있음을보여준다.
    번역하기

    Rotary positional embedding (RoPE)은현대의대규모 언어 모델에서널리 사용되고있으나, 최근연구 들은RoPE가attention sink나 중간문맥 감도저하와같은장문 문맥 추론을방해하는 inductive bias를 유발할 수...

    Rotary positional embedding (RoPE)은현대의대규모 언어 모델에서널리 사용되고있으나, 최근연구
    들은RoPE가attention sink나 중간문맥 감도저하와같은장문 문맥 추론을방해하는 inductive bias를
    유발할 수있음을시사한다. 본연구에서는 추가적인학습없이, 사전학습된모델내 일부 attention head
    를 No Positional Encoding (NoPE)으로변환함으로써 RoPE를 부분적으로제거할 수있는지를 분석한다.
    Llama-3.2-3B-Instruct 모델을대상으로RULER 벤치마크에서4K–8K 문맥 길이범위에 걸쳐 head- 및
    layer-단위의체계적인ablation 실험을수행한 결과, NoPE 변환이항상 성능 저하를 초래하지는 않음을
    확인하였다. 초기 레이어는 RoPE 제거에 매우민감한 반면, 후반레이어에는 NoPE로대체했을때 일관된
    성능 향상을보이는 attention head들이존재한다. 또한 position-awareness-guided 선택과무작위head 선택
    모두에서성능 향상이관찰되어, 이러한 이점이정확한 head 식별보다는 positional bias를 완화하는 데서
    비롯됨을시사한다.
    마지막으로attention 분석을통해 NoPE 변환이attention을중간문맥 토큰으로재분배함으로써 attention
    sink 현상을완화함을보인다. 이러한 결과는 재학습없이도inference 단계에서선택적으로NoPE를 적용
    하는 것이사전학습된모델의장문 문맥 처리 능력을향상시킬 수있음을보여준다.

    더보기

    목차 (Table of Contents)

    • 1 Abstract 1
    • 2 Introduction 3
    • 3 Related Work 4
    • 1 Abstract 1
    • 2 Introduction 3
    • 3 Related Work 4
    • 3.1 Positional Encoding and Rotary Embedding 4
    • 3.2 Long-Context Awareness 4
    • 3.3 RoPE Variations 5
    • 3.4 Head Specialization in Multi-Head Attention 5
    • 4 Method 6
    • 4.1 Manual Head Selection 6
    • 4.1.1 Position-Awareness-Guided Selection 6
    • 4.1.2 Random Selection 7
    • 5 Experiment 7
    • 5.1 Baseline Configuration 7
    • 5.2 Number of NoPE Heads per Layer 7
    • 5.3 Number of NoPE Layers 9
    • 5.4 Position-Awareness-Guided vs. Random Head Selection 10
    • 5.5 Why Does NoPE Replacement Improve Long-Context Performance? 11
    • 6 Conclusion 13
    • A Additional Results on Qwen-4B 14
    • B Abstract in Korean 15
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼