RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    정적 분할 알고리즘을 이용한 에너지 효율적 비전 트랜스포머 하드웨어 가속기 설계 = An Energy-Efficient Vision Transformer Hardware Accelerator Using Static Partitioning Algorithm

    한글로보기

    https://www.riss.kr/link?id=T17198100

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 논문에서는 정적 분할 알고리즘 기법을 이용한 에너지 효율적 비전 트랜스포머(Vision Transformer) 하드웨어 가속기 설계를 제안한다. Vision Transformer의 경우 컴퓨터 비전 분야에서 다양하게 응용되고 있는 Transformer 계열의 알고리즘이다. 하지만 파라미터가 너무 크고 에너지 소모량이 많아 엣지 디바이스 수준에서 사용하는 데 어려움이 있기 때문에 최적화가 필수이다. ViT 알고리즘 내부의 어텐션 매커니즘(Attention Mechanism)은 연산이 복잡하며 여러 번 반복해야 하므로 실행시간이 긴 연산이다. Array의 크기 자체가 크기 때문에 이미 메모리 소모가 많은 데다 Matrix Multiply를 하는 연산이기 때문에 최적 설계 되어야 할 핵심 모듈이 다. 제안하는 ViT 가속기는 Attention Mechanism의 Sparse Pattern을 최 적화하여 활용하기 위한 Sparse Matrix Processor(SMP)와 Multi Layer Perceptron을 효율적으로 계산하기 위해 Systolic Array 구조를 이용한 Dense Matrix Processor(DMP)로 구성되어있다. SMP의 경우, Static Partitioning Algorithm을 활용하여 Sparse Matrix Multiply 을 최적화하 기 위한 구조로 설계했다. Sparse Matrix 자체의 효율을 높이기 위해서 BCSR(Block Compressed Sparse Row) Format을 활용하였다. DMP의 경우, 내부구조는 Systolic Array 구조로 설계하였고 다수의 DMP을 활용 해 큰 행렬 연산 시에 데이터 처리량을 높였다. 제안된 에너지 효율적인 ViT 가속기는 Xilinx Vitis HLS, Vivado 툴과 FPGA ZYNQ ZCU-102보드를 이용하여 구현하였으며, 성능 비교 및 검증 을 하였다. FPGA Utilization 측정 결과, LUT Slice는 163K(59.4%), DSP 블록은 1,452(57.6%), BRAM은 1,049(57.5%)가 사용되었다. 클럭 속도는 300MHz로 동작한다. 제안된 ViT 가속기의 데이터 처리율은 107.4 GOP/s(Giga Operation per Second)이며, 에너지 효율은 10.89 GOP/J로 비슷한 크기의 ViT 가속기들 대비 최대 225% 향상된 데이터 처리율을 갖 는다. 본 논문에서 제안하는 정적 분할 알고리즘 기법을 이용한 에너지 효율적 비전 트랜스포머 하드웨어 가속기는 저전력 구동이 요구되는 딥러닝 기반의 다양한 어플리케이션에서 높은 성능을 발휘할 것으로 기대된다. Key Words: 정적 분할 알고리즘(Static Partitioning Algorithm), 비전 트랜스포머(Vision Transformer), 하드웨어 가속기(Hardware Accelerator), 컴퓨터 비전(Computer Vision), 어텐션 매커니즘 (Attention Mechanism), Systolic Array, BCSR(Block Compressed Sparse Row), FPGA(Field Programmable Gate Array)
    번역하기

    본 논문에서는 정적 분할 알고리즘 기법을 이용한 에너지 효율적 비전 트랜스포머(Vision Transformer) 하드웨어 가속기 설계를 제안한다. Vision Transformer의 경우 컴퓨터 비전 분야에서 다양하게 ...

    본 논문에서는 정적 분할 알고리즘 기법을 이용한 에너지 효율적 비전 트랜스포머(Vision Transformer) 하드웨어 가속기 설계를 제안한다. Vision Transformer의 경우 컴퓨터 비전 분야에서 다양하게 응용되고 있는 Transformer 계열의 알고리즘이다. 하지만 파라미터가 너무 크고 에너지 소모량이 많아 엣지 디바이스 수준에서 사용하는 데 어려움이 있기 때문에 최적화가 필수이다. ViT 알고리즘 내부의 어텐션 매커니즘(Attention Mechanism)은 연산이 복잡하며 여러 번 반복해야 하므로 실행시간이 긴 연산이다. Array의 크기 자체가 크기 때문에 이미 메모리 소모가 많은 데다 Matrix Multiply를 하는 연산이기 때문에 최적 설계 되어야 할 핵심 모듈이 다. 제안하는 ViT 가속기는 Attention Mechanism의 Sparse Pattern을 최 적화하여 활용하기 위한 Sparse Matrix Processor(SMP)와 Multi Layer Perceptron을 효율적으로 계산하기 위해 Systolic Array 구조를 이용한 Dense Matrix Processor(DMP)로 구성되어있다. SMP의 경우, Static Partitioning Algorithm을 활용하여 Sparse Matrix Multiply 을 최적화하 기 위한 구조로 설계했다. Sparse Matrix 자체의 효율을 높이기 위해서 BCSR(Block Compressed Sparse Row) Format을 활용하였다. DMP의 경우, 내부구조는 Systolic Array 구조로 설계하였고 다수의 DMP을 활용 해 큰 행렬 연산 시에 데이터 처리량을 높였다. 제안된 에너지 효율적인 ViT 가속기는 Xilinx Vitis HLS, Vivado 툴과 FPGA ZYNQ ZCU-102보드를 이용하여 구현하였으며, 성능 비교 및 검증 을 하였다. FPGA Utilization 측정 결과, LUT Slice는 163K(59.4%), DSP 블록은 1,452(57.6%), BRAM은 1,049(57.5%)가 사용되었다. 클럭 속도는 300MHz로 동작한다. 제안된 ViT 가속기의 데이터 처리율은 107.4 GOP/s(Giga Operation per Second)이며, 에너지 효율은 10.89 GOP/J로 비슷한 크기의 ViT 가속기들 대비 최대 225% 향상된 데이터 처리율을 갖 는다. 본 논문에서 제안하는 정적 분할 알고리즘 기법을 이용한 에너지 효율적 비전 트랜스포머 하드웨어 가속기는 저전력 구동이 요구되는 딥러닝 기반의 다양한 어플리케이션에서 높은 성능을 발휘할 것으로 기대된다. Key Words: 정적 분할 알고리즘(Static Partitioning Algorithm), 비전 트랜스포머(Vision Transformer), 하드웨어 가속기(Hardware Accelerator), 컴퓨터 비전(Computer Vision), 어텐션 매커니즘 (Attention Mechanism), Systolic Array, BCSR(Block Compressed Sparse Row), FPGA(Field Programmable Gate Array)

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In this paper, we propose the design of an energy-efficient Vision Transformer hardware accelerator using static partitioning algorithm techniques. Vision Transformer is a Transformer-based algorithm widely applied in computer vision fields. However, optimization is essential as it is difficult to implement on edge devices due to its large parameters and high energy consumption. The Attention Mechanism within the ViT algorithm is a time-consuming operation that requires complex calculations and multiple iterations. The Attention module needs to be optimally designed as it already consumes significant memory due to its large array size and involves matrix multiplication operations. The proposed ViT accelerator consists of a Sparse Matrix Processor (SMP) for optimizing and utilizing the Sparse Pattern of the Attention Mechanism, and a Dense Matrix Processor (DMP) using Systolic Array architecture for efficient Multi Layer Perceptron calculations. The SMP is designed with a structure that optimizes Sparse Matrix Multiply using Static Partitioning Algorithm. To improve the efficiency of the Sparse Matrix itself, BCSR (Block Compressed Sparse Row) Format was utilized. For the DMP, the internal structure was designed with Systolic Array architecture, and multiple DMPs were utilized to increase data throughput during large matrix operations. The proposed energy-efficient ViT accelerator was implemented using Xilinx Vitis HLS, Vivado tools, and FPGA ZYNQ ZCU-102 board, and its performance was compared and verified. According to FPGA Utilization measurements, LUT Slice used 163K (59.4%), DSP blocks used 1,452 (57.6%), and BRAM used 1,049 (57.5%). The clock speed operates at 300MHz. The proposed ViT accelerator has a data throughput of 107.4 GOP/s (Giga Operation per Second) and an energy efficiency of 10.89 GOP/J, achieving up to 225% improved data throughput compared to ViT accelerators of similar size. The energy-efficient Vision Transformer hardware accelerator using static partitioning algorithm techniques proposed in this paper is expected to deliver high performance in various deep learning-based applications requiring low-power operation. Key Words: Static Partitioning Algorithm, Vision Transformer, Hardware Accelerator, Computer Vision, Attention Mechanism, Systolic Array, BCSR(Block Compressed Sparse Row), FPGA(Field Programmable Gate Array)
    번역하기

    In this paper, we propose the design of an energy-efficient Vision Transformer hardware accelerator using static partitioning algorithm techniques. Vision Transformer is a Transformer-based algorithm widely applied in computer vision fields. However, ...

    In this paper, we propose the design of an energy-efficient Vision Transformer hardware accelerator using static partitioning algorithm techniques. Vision Transformer is a Transformer-based algorithm widely applied in computer vision fields. However, optimization is essential as it is difficult to implement on edge devices due to its large parameters and high energy consumption. The Attention Mechanism within the ViT algorithm is a time-consuming operation that requires complex calculations and multiple iterations. The Attention module needs to be optimally designed as it already consumes significant memory due to its large array size and involves matrix multiplication operations. The proposed ViT accelerator consists of a Sparse Matrix Processor (SMP) for optimizing and utilizing the Sparse Pattern of the Attention Mechanism, and a Dense Matrix Processor (DMP) using Systolic Array architecture for efficient Multi Layer Perceptron calculations. The SMP is designed with a structure that optimizes Sparse Matrix Multiply using Static Partitioning Algorithm. To improve the efficiency of the Sparse Matrix itself, BCSR (Block Compressed Sparse Row) Format was utilized. For the DMP, the internal structure was designed with Systolic Array architecture, and multiple DMPs were utilized to increase data throughput during large matrix operations. The proposed energy-efficient ViT accelerator was implemented using Xilinx Vitis HLS, Vivado tools, and FPGA ZYNQ ZCU-102 board, and its performance was compared and verified. According to FPGA Utilization measurements, LUT Slice used 163K (59.4%), DSP blocks used 1,452 (57.6%), and BRAM used 1,049 (57.5%). The clock speed operates at 300MHz. The proposed ViT accelerator has a data throughput of 107.4 GOP/s (Giga Operation per Second) and an energy efficiency of 10.89 GOP/J, achieving up to 225% improved data throughput compared to ViT accelerators of similar size. The energy-efficient Vision Transformer hardware accelerator using static partitioning algorithm techniques proposed in this paper is expected to deliver high performance in various deep learning-based applications requiring low-power operation. Key Words: Static Partitioning Algorithm, Vision Transformer, Hardware Accelerator, Computer Vision, Attention Mechanism, Systolic Array, BCSR(Block Compressed Sparse Row), FPGA(Field Programmable Gate Array)

    더보기

    목차 (Table of Contents)

    • 제 1 장 서론 1
    • 제 2 장 본론 4
    • 제 1 절 신경망 모델 4
    • 1.1. 심층 신경망 (Deep Neural Networks) 4
    • 제 1 장 서론 1
    • 제 2 장 본론 4
    • 제 1 절 신경망 모델 4
    • 1.1. 심층 신경망 (Deep Neural Networks) 4
    • 1.2. 합성곱 신경망 (Convolutional Neural Networks) 8
    • 1.3. 순환 신경망 (Recurrent Neural Networks) 11
    • 1.4. 어텐션 (Attention) 13
    • 1.5. 트랜스포머 (Transformer) 15
    • 제 2 절 Vision Transformer 19
    • 2.1. Vision Transformer의 개념 19
    • 2.2. Vision Transformer 계열의 주요 모델 21
    • 2.3. 다양한 딥러닝 하드웨어 가속기 소개 24
    • 제 3 절 비전 트랜스포머 하드웨어 가속기 구조 30
    • 3.1. 비전 트랜스포머 가속기를 위한 정적 할당 알고리즘 31
    • 3.2. 제안하는 정적 분할 알고리즘을 이용한 에너지 효율적 비전 트랜스포머 하드웨어 가속기 구조 37
    • 3.3. 비선형 활성화 함수 하드웨어 최적화 40
    • 제 4 절 구현 및 성능 비교 43
    • 4.1. 제안하는 정적 분할 알고리즘을 이용한 에너지 효율적 비전 트랜스포머 하드웨어 가속기의 기능 및 성능 분석 45
    • 4.2. 제안하는 정적 분할 알고리즘을 이용한 에너지 효율적 비전 트랜스포머 하드웨어 가속기의 구현 결과 분석 50
    • 제 3 장 결론 51
    • 참고문헌 53
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼