RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Accelerating Memory-Bound Operations in Emerging Deep Neural Networks = 행렬곱 비중이 낮은 심층신경망에서 메모리-바운드 연산 가속화

    한글로보기

    https://www.riss.kr/link?id=T17449762

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Deep neural networks (DNNs) have achieved remarkable success across a wide range of applications. However, their massive computational requirements have motivated the exploration of more compute-efficient alternative operations. Many of these emerging operations, while reducing arithmetic workload, exhibit low arithmetic intensity and thus tend to be memory-bound on modern hardware. As the attainable memory bandwidth lags far behind the available compute throughput, these algorithms increasingly suffer from the memory wall, which significantly hinders their performance.

    This thesis presents acceleration techniques for two emerging memory-bound operations in DNN workloads, both leveraging kernel fusion. The first target is the tensor-product operation used in graph neural networks for predicting atomic energies in molecular dynamics simulations. To accelerate this memory-bound computation, we employ a fusion strategy that combines multiple kernels while exploiting computational sparsity and input reuse. The second target is the state space model-based convolution operation, an emerging alternative to self-attention in large language models. For this case, we design a custom accelerator architecture that supports real-time operand decompression, thereby reducing the required on-chip SRAM capacity for the fused kernel and improving the overall performance and energy efficiency.
    번역하기

    Deep neural networks (DNNs) have achieved remarkable success across a wide range of applications. However, their massive computational requirements have motivated the exploration of more compute-efficient alternative operations. Many of these emerging...

    Deep neural networks (DNNs) have achieved remarkable success across a wide range of applications. However, their massive computational requirements have motivated the exploration of more compute-efficient alternative operations. Many of these emerging operations, while reducing arithmetic workload, exhibit low arithmetic intensity and thus tend to be memory-bound on modern hardware. As the attainable memory bandwidth lags far behind the available compute throughput, these algorithms increasingly suffer from the memory wall, which significantly hinders their performance.

    This thesis presents acceleration techniques for two emerging memory-bound operations in DNN workloads, both leveraging kernel fusion. The first target is the tensor-product operation used in graph neural networks for predicting atomic energies in molecular dynamics simulations. To accelerate this memory-bound computation, we employ a fusion strategy that combines multiple kernels while exploiting computational sparsity and input reuse. The second target is the state space model-based convolution operation, an emerging alternative to self-attention in large language models. For this case, we design a custom accelerator architecture that supports real-time operand decompression, thereby reducing the required on-chip SRAM capacity for the fused kernel and improving the overall performance and energy efficiency.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    심층신경망의 혁신적인 발전은 다양한 응용 분야에서 놀라운 성과를 이끌어냈지만, 학습 및 추론 과정에서 요구되는 막대한 연산량으로 인해, 계산 효율이 높은 대체 연산 방식에 대한 연구가 활발히 이루어지고 있는 추세이다. 그러나 계산 효율적인 연산들은 산술 집약도가 낮아 메모리 병목이 발생하기 쉽고, 현재 하드웨어에서 메모리 대역폭의 성능이 연산 속도에 비해 크게 낮기 때문에 결과적으로 성능 저하가 불가피한 상황이다.

    본 논문은 DNN 워크로드에서 새롭게 등장한 행렬곱 비중이 낮은 두 가지 연산의 메모리 병목을 완화하기 위해, 커널 융합을 활용한 가속화 기법을 제안한다. 첫 번째 대상은 분자동역학 시뮬레이션에서 원자 에너지를 예측하기 위해 등변적 그래프 신경망에서 사용되는 텐서곱 연산이다. 이 연산의 메모리 병목을 해소하기 위해, 연산 내 희소성과 피연산자 재사용을 활용하는 융합 커널인 FlashTP를 제안하였다. 두 번째 대상은 대규모 언어 모델에서 셀프 어텐션을 대체하는 새로운 방식으로 주목받고 있는 상태공간모델 기반 합성곱 연산이다. 이 경우, 융합된 커널의 실행에 필요한 칩 내부 SRAM 용량을 줄이고 전반적인 성능 및 에너지 효율을 향상시키기 위해, 실시간 피연산자 압축 해제를 지원하는 맞춤형 가속기 아키텍처인 VGA를 설계하였다.
    번역하기

    심층신경망의 혁신적인 발전은 다양한 응용 분야에서 놀라운 성과를 이끌어냈지만, 학습 및 추론 과정에서 요구되는 막대한 연산량으로 인해, 계산 효율이 높은 대체 연산 방식에 대한 연구...

    심층신경망의 혁신적인 발전은 다양한 응용 분야에서 놀라운 성과를 이끌어냈지만, 학습 및 추론 과정에서 요구되는 막대한 연산량으로 인해, 계산 효율이 높은 대체 연산 방식에 대한 연구가 활발히 이루어지고 있는 추세이다. 그러나 계산 효율적인 연산들은 산술 집약도가 낮아 메모리 병목이 발생하기 쉽고, 현재 하드웨어에서 메모리 대역폭의 성능이 연산 속도에 비해 크게 낮기 때문에 결과적으로 성능 저하가 불가피한 상황이다.

    본 논문은 DNN 워크로드에서 새롭게 등장한 행렬곱 비중이 낮은 두 가지 연산의 메모리 병목을 완화하기 위해, 커널 융합을 활용한 가속화 기법을 제안한다. 첫 번째 대상은 분자동역학 시뮬레이션에서 원자 에너지를 예측하기 위해 등변적 그래프 신경망에서 사용되는 텐서곱 연산이다. 이 연산의 메모리 병목을 해소하기 위해, 연산 내 희소성과 피연산자 재사용을 활용하는 융합 커널인 FlashTP를 제안하였다. 두 번째 대상은 대규모 언어 모델에서 셀프 어텐션을 대체하는 새로운 방식으로 주목받고 있는 상태공간모델 기반 합성곱 연산이다. 이 경우, 융합된 커널의 실행에 필요한 칩 내부 SRAM 용량을 줄이고 전반적인 성능 및 에너지 효율을 향상시키기 위해, 실시간 피연산자 압축 해제를 지원하는 맞춤형 가속기 아키텍처인 VGA를 설계하였다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • List of Figures v
    • Abstract i
    • Contents ii
    • List of Figures v
    • List of Tables x
    • Chapter 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Thesis Overview 2
    • 1.3 Bibliographic Remarks 3
    • Chapter 2 Background 4
    • 2.1 Memory Wall in Emerging DNN Algorithms 4
    • 2.2 Software Solution for Memory-Bound Workloads 6
    • 2.3 Hardware Solution for Memory-Bound Workloads 8
    • Chapter 3 FlashTP: Fused, Sparsity-Aware Tensor Product for Machine Learning Interatomic Potentials 11
    • 3.1 Introduction 11
    • 3.2 Background 13
    • 3.2.1 Machine Learning Interatomic Potential (MLIP) 13
    • 3.2.2 Equivariant MLIP 14
    • 3.2.3 Tensor-Product Layer in Equivariant MLIP 18
    • 3.3 Inefficiency in Tensor-Product Layer 19
    • 3.3.1 Memory Traffic from Intermediate Data 19
    • 3.3.2 Peak Memory Spikes due to Output Data 21
    • 3.3.3 Ineffectual Computation due to Sparsity 22
    • 3.4 FlashTP 22
    • 3.4.1 Overview 22
    • 3.4.2 Kernel Fusion 23
    • 3.4.3 Applying Sparsity in Tensor-Product 26
    • 3.4.4 Path-Aggregation 28
    • 3.5 Implementation 29
    • 3.6 Evaluation 30
    • 3.6.1 Methodology 30
    • 3.6.2 Kernel Microbenchmark 34
    • 3.6.3 End-to-end Speedup 38
    • 3.7 Related Work 40
    • 3.8 Conclusion 41
    • Chapter 4 VGA: Hardware Accelerator for Scalable Long Sequence Model Inference 42
    • 4.1 Introduction 42
    • 4.2 Background 46
    • 4.2.1 Limitations of Self Attention 46
    • 4.2.2 Global Convolution Models 46
    • 4.2.3 Cooley-Tukey FFT Algorithm 48
    • 4.2.4 State Space Model (SSM)-based Global Convolution 50
    • 4.2.5 State Passing 51
    • 4.3 Analysis of H3 Computation 53
    • 4.3.1 H3 Block Runtime Breakdown 53
    • 4.3.2 ROI Breakdown 55
    • 4.3.3 Custom Accelerator Solution 56
    • 4.4 Architecture Design 57
    • 4.4.1 Overview 57
    • 4.4.2 Computational Components 58
    • 4.4.3 Operation Mapping 60
    • 4.4.4 Other Hardware Components 63
    • 4.4.5 System-level Issues 64
    • 4.5 Evaluation 66
    • 4.5.1 Methodology 66
    • 4.5.2 ROI Speedups 69
    • 4.5.3 Model Speedups 70
    • 4.5.4 Source of Efficiency 71
    • 4.5.5 Sensitivity Studies 73
    • 4.5.6 Area/Power Analysis 74
    • 4.5.7 Discussion 74
    • 4.6 Related Work 76
    • 4.7 Conclusion 77
    • Chapter 5 Conclusion and Future Directions 79
    • 5.1 Summary of Contributions 79
    • 5.2 Future Directions 80
    • Bibliography 82
    • 국문초록 100
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼