RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    An Optimized Collective Communication Framework for PIM-Enabled Systems = 범용 PIM 기반 시스템을 위한 효율적인 집단 통신 라이브러리

    한글로보기

    https://www.riss.kr/link?id=T17450742

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent dual in-line memory modules (DIMMs) increasingly support processing-in-memory (PIM) by integrating processing elements (PEs) within memory banks, enabling applications to mitigate the data movement bottleneck. While many highly parallel applications benefit from PIM-enabled DIMMs, performance gains are often constrained by substantial inter-PE collective communication overhead, primarily due to slow CPU-mediated communication methods. Although prior work has attempted to address this bottleneck, existing solutions lack the flexibility and performance necessary for diverse applications.
    This dissertation presents PID-Comm, a fast and flexible collective communication framework for commodity PIM-enabled DIMMs. PID-Comm introduces a multi-dimensional hypercube abstraction for PE organization, enabling concurrent collective communication among PEs within specific hypercube dimensions. Building on this abstraction, PID-Comm provides high-performance implementations of eight inter-PE collective communication patterns optimized for the DIMMs. Evaluation on 16 UPMEM DIMMs using representative parallel algorithms demonstrates that PID-Comm achieves up to 5.19× performance improvement over existing implementations.
    번역하기

    Recent dual in-line memory modules (DIMMs) increasingly support processing-in-memory (PIM) by integrating processing elements (PEs) within memory banks, enabling applications to mitigate the data movement bottleneck. While many highly parallel applica...

    Recent dual in-line memory modules (DIMMs) increasingly support processing-in-memory (PIM) by integrating processing elements (PEs) within memory banks, enabling applications to mitigate the data movement bottleneck. While many highly parallel applications benefit from PIM-enabled DIMMs, performance gains are often constrained by substantial inter-PE collective communication overhead, primarily due to slow CPU-mediated communication methods. Although prior work has attempted to address this bottleneck, existing solutions lack the flexibility and performance necessary for diverse applications.
    This dissertation presents PID-Comm, a fast and flexible collective communication framework for commodity PIM-enabled DIMMs. PID-Comm introduces a multi-dimensional hypercube abstraction for PE organization, enabling concurrent collective communication among PEs within specific hypercube dimensions. Building on this abstraction, PID-Comm provides high-performance implementations of eight inter-PE collective communication patterns optimized for the DIMMs. Evaluation on 16 UPMEM DIMMs using representative parallel algorithms demonstrates that PID-Comm achieves up to 5.19× performance improvement over existing implementations.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 메모리 모듈(DIMM)은 메모리 뱅크 내에 연산 요소를 통합하여 메모리 내 연산을 지원함으로써 데이터 이동 병목 현상을 완화할 수 있게 한다. 많은 병렬 애플리케이션들이 PIM 지원 DIMM의 이점을 누리고 있지만, PE 간의 비효율적인 통신 방식으로 인해 성능 향상에는 제약이 있다. 기존 연구들은 이러한 통신 병목 현상을 해결하고자 시도했지만, 여전히 다향한 애플리케이션에 필요한 유연성과 성능은 부족하다.
    본 연구에서는 상용 PIM 지원 DIMM을 위한 빠르고 유연한 집합 통신 프레임워크인 PID-Comm 을 제시한다. PID-Comm은 PE 조직화를 위한 다차원 하이퍼큐브 추상화를 도입하여, 특정 하이퍼큐브 차원 내에서 PE 간 동시 집합 통신을 가능하게 한다. 이러한 추상화를 기반으로, PID-Comm은 DIMM에 최적화된 8가지 집합 통신 패턴의 고성능 구현을 제공한다. 대표적인 병렬 알고리즘을 사용한 평가 결과, PID-Comm은 기존 구현 대비 최대 5.19배의 성능 향상을 달성하였다.
    번역하기

    최근 메모리 모듈(DIMM)은 메모리 뱅크 내에 연산 요소를 통합하여 메모리 내 연산을 지원함으로써 데이터 이동 병목 현상을 완화할 수 있게 한다. 많은 병렬 애플리케이션들이 PIM 지원 DIMM의 ...

    최근 메모리 모듈(DIMM)은 메모리 뱅크 내에 연산 요소를 통합하여 메모리 내 연산을 지원함으로써 데이터 이동 병목 현상을 완화할 수 있게 한다. 많은 병렬 애플리케이션들이 PIM 지원 DIMM의 이점을 누리고 있지만, PE 간의 비효율적인 통신 방식으로 인해 성능 향상에는 제약이 있다. 기존 연구들은 이러한 통신 병목 현상을 해결하고자 시도했지만, 여전히 다향한 애플리케이션에 필요한 유연성과 성능은 부족하다.
    본 연구에서는 상용 PIM 지원 DIMM을 위한 빠르고 유연한 집합 통신 프레임워크인 PID-Comm 을 제시한다. PID-Comm은 PE 조직화를 위한 다차원 하이퍼큐브 추상화를 도입하여, 특정 하이퍼큐브 차원 내에서 PE 간 동시 집합 통신을 가능하게 한다. 이러한 추상화를 기반으로, PID-Comm은 DIMM에 최적화된 8가지 집합 통신 패턴의 고성능 구현을 제공한다. 대표적인 병렬 알고리즘을 사용한 평가 결과, PID-Comm은 기존 구현 대비 최대 5.19배의 성능 향상을 달성하였다.

    더보기

    목차 (Table of Contents)

    • Chapter 1: Introduction 1
    • Chapter 2: Background 4
    • 2.1 PIM-enabled DIMMs and Entangled Groups 4
    • 2.2 Domain Transfer 5
    • 2.2.1 Collective Communications 6
    • Chapter 1: Introduction 1
    • Chapter 2: Background 4
    • 2.1 PIM-enabled DIMMs and Entangled Groups 4
    • 2.2 Domain Transfer 5
    • 2.2.1 Collective Communications 6
    • Chapter 3: Motivation 8
    • 3.1 Conventional Communication Models and Libraries 8
    • 3.2 Lack of a Flexible Communication Model 10
    • Chapter 4: PID-Comm Communication Model 12
    • 4.1 Design Goals and Challenges 12
    • 4.2 Virtual Hypercube Communication Model 13
    • 4.2.1 User-defined hypercube configuration 13
    • 4.2.2 Cube slices as communication groups 13
    • 4.2.3 Multi-instance invocation 14
    • 4.3 Mapping Virtual Hypercube to Physical PEs 14
    • Chapter 5: PID-Comm Library 17
    • 5.1 PID-Comm Performance Optimization Techniques 19
    • 5.1.1 PE-assisted reordering 19
    • 5.1.2 In-register modulation 20
    • 5.1.3 Cross-domain modulation 20
    • 5.2 Other Collective Primitives 20
    • 5.2.1 AllGather (AG) 22
    • 5.2.2 ReduceScatter (RS) 22
    • 5.2.3 AllReduce (AR) 23
    • 5.2.4 Primitives with Roots 23
    • 5.3 Extension to General Cases 24
    • Chapter 6: PID-Comm Programming Framework 26
    • 6.1 Communication Framework 26
    • 6.2 Implementation 27
    • Chapter 7: Benchmark Applications 29
    • 7.1 Deep Learning Recommendation Model 30
    • 7.2 Graph Neural Networks 31
    • 7.3 Breadth-First Search 32
    • 7.4 Connected Components 32
    • 7.5 Multi-Layer Perceptron 32
    • Chapter 8: Evaluation 34
    • 8.1 Experimental Setup 34
    • 8.2 Performance of Supported Primitives 35
    • 8.3 Performance of Benchmark Applications 35
    • 8.4 Ablation Study 37
    • 8.5 Sensitivity Study 39
    • 8.6 Sensitivity Study on Different Word Bits 41
    • 8.7 Comparison to CPU-only Systems 42
    • 8.8 Comparison to Other Hierarchy-aware Approaches 44
    • Chapter 9: discussion 46
    • 9.1 PID-Comm on Other PIM Architectures/Systems 46
    • 9.2 Hardward Implications 47
    • Chapter 10: Related Work
    • 10.1 Processing-in-Memory 50
    • 10.2 Collective Communication 51
    • Chapter 11: Conclusion 52
    • 초록 64
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼