RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    On the head redundancy in Swin transformer for image classification

    한글로보기

    https://www.riss.kr/link?id=T16373270

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    Transformer 모델은 원래 자연어 처리를 위해 고안된 모델이지만, computer vision 등의 타 분야에서도 Transformer를 활용하려는 연구가 매우 활발하게 진행되고 있다. 자연어 처리 분야에서 이용되는 Transformer 기반 모델의 Multi-Head Self-Attention (MHSA) 모듈을 구성하는 각 헤드간의 중복성에 관한 연구는 존재하지만, computer vision 분야에서 이용되는 Transformer 기반 모델에 대해 이를 연구한 바는 없다. 본 논문에서는 Swin Transformer 모델의 MHSA 모듈에서 각 헤드를 제거했을 때 ImageNet 데이터셋의 이미지 분류 작업에서 모델 성능의 변화를 측정하였고, 이를 통해 몇몇 헤드는 다른 헤드와 중복성이 존재하여, 그 헤드를 제거하더라도 정확도의 하락이 적거나 오히려 상승함을 밝혔다. 또한 stage 3을 구성하는 헤드 중 절반을 제거하여 비교적 적은 정확도 감소를 대가로 모델을 경량화할 수 있음을 보였다. 마지막으로, 헤드의 중복성에 영향을 미칠 것으로 예상되는 3개의 잠재요소로 출력 행렬의 노름의 평균값, 입력에 따른 출력 행렬 간의 불변성, 각 헤드 간 attention map의 코사인 유사도를 제시하였고, 그 중 전자 2개의 요소와 모델 성능 간에 상관관계가 존재함을 발견하였다.
    번역하기

    Transformer 모델은 원래 자연어 처리를 위해 고안된 모델이지만, computer vision 등의 타 분야에서도 Transformer를 활용하려는 연구가 매우 활발하게 진행되고 있다. 자연어 처리 분야에서 이용되는 ...

    Transformer 모델은 원래 자연어 처리를 위해 고안된 모델이지만, computer vision 등의 타 분야에서도 Transformer를 활용하려는 연구가 매우 활발하게 진행되고 있다. 자연어 처리 분야에서 이용되는 Transformer 기반 모델의 Multi-Head Self-Attention (MHSA) 모듈을 구성하는 각 헤드간의 중복성에 관한 연구는 존재하지만, computer vision 분야에서 이용되는 Transformer 기반 모델에 대해 이를 연구한 바는 없다. 본 논문에서는 Swin Transformer 모델의 MHSA 모듈에서 각 헤드를 제거했을 때 ImageNet 데이터셋의 이미지 분류 작업에서 모델 성능의 변화를 측정하였고, 이를 통해 몇몇 헤드는 다른 헤드와 중복성이 존재하여, 그 헤드를 제거하더라도 정확도의 하락이 적거나 오히려 상승함을 밝혔다. 또한 stage 3을 구성하는 헤드 중 절반을 제거하여 비교적 적은 정확도 감소를 대가로 모델을 경량화할 수 있음을 보였다. 마지막으로, 헤드의 중복성에 영향을 미칠 것으로 예상되는 3개의 잠재요소로 출력 행렬의 노름의 평균값, 입력에 따른 출력 행렬 간의 불변성, 각 헤드 간 attention map의 코사인 유사도를 제시하였고, 그 중 전자 2개의 요소와 모델 성능 간에 상관관계가 존재함을 발견하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Although the Transformer model was originally designed for natural language processing (NLP), many studies are being actively conducted in utilizing the Transformer in other fields such as computer vision. There is a study on the redundancy between attention heads in multi-head self-attention (MHSA) modules of Transformer-based models in NLP, but there is no study on Transformer-based models in computer vision. In this thesis, we measure the change of model performance on the ImageNet image classification task when each head in MHSA layers in Swin Transformer is removed. Several heads are redundant, so the model accuracy slightly decreases or rather increases. In addition, it is shown that the model can be compressed by removing half of the heads in stage 3, in exchange for insignificant accuracy loss. Finally, we offer three factors that are expected to affect the head redundancy; the mean of norms of output matrices for various inputs, the invariance of output matrices for various inputs, and the cosine similarity between attention maps of each head. It is proved that the former two factors are correlated with head redundancy.
    번역하기

    Although the Transformer model was originally designed for natural language processing (NLP), many studies are being actively conducted in utilizing the Transformer in other fields such as computer vision. There is a study on the redundancy between at...

    Although the Transformer model was originally designed for natural language processing (NLP), many studies are being actively conducted in utilizing the Transformer in other fields such as computer vision. There is a study on the redundancy between attention heads in multi-head self-attention (MHSA) modules of Transformer-based models in NLP, but there is no study on Transformer-based models in computer vision. In this thesis, we measure the change of model performance on the ImageNet image classification task when each head in MHSA layers in Swin Transformer is removed. Several heads are redundant, so the model accuracy slightly decreases or rather increases. In addition, it is shown that the model can be compressed by removing half of the heads in stage 3, in exchange for insignificant accuracy loss. Finally, we offer three factors that are expected to affect the head redundancy; the mean of norms of output matrices for various inputs, the invariance of output matrices for various inputs, and the cosine similarity between attention maps of each head. It is proved that the former two factors are correlated with head redundancy.

    더보기

    목차 (Table of Contents)

    • List of Figures
    • Abstract
    • Chapter 1. Introduction
    • Chapter 2. Backgrounds and Previous Works
    • List of Figures
    • Abstract
    • Chapter 1. Introduction
    • Chapter 2. Backgrounds and Previous Works
    • 2.1. Self-Attention Layer
    • 2.2. Applications of Self-Attention in Vision
    • 2.3. Swin Transformer
    • 2.4. Pruning Filters for Efficient ConvNets
    • Chapter 3. Investigation of the Head Redundancy
    • 3.1. Experimental Setup
    • 3.2. Removing One Head
    • 3.3. Removing Multiple Heads Iteratively
    • Chapter 4. Factors Affecting the Head Redundancy
    • 4.1. Mean of Norms of Output Matrices for Various Inputs
    • 4.2. Invariance of Output Matrices for Various Inputs
    • 4.3. Cosine Similarity Rate Between Attention Maps of Each Head
    • 4.4. Correlations & Multiple Linear Regression
    • Chapter 5. Conclusion
    • References
    • Abstract in Korean
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼