RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Proactive congestion avoidance for distributed deep learning

    한글로보기

    https://www.riss.kr/link?id=T15783368

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This thesis presents “Proactive Congestion Notification” (PCN), a congestionavoidance technique for distributed deep learning (DDL). DDL is widely used to scale out and accelerate deep neural network training. In DDL, each worker trains a copy of the deep learning model with different training inputs and synchronizes the model gradients at the end of each iteration. However, it is well known that the network communication for synchronizing model parameters is the main bottleneck in DDL. Our key observation is that the DDL architecture makes each worker generate burst traffic every iteration, which causes network congestion and in turn degrades the throughput of DDL traffic. Based on this observation, the key idea behind PCN is to prevent potential congestion by proactively regulating the switch queue length before DDL burst traffic arrives at the switch, which prepares the switches for handling incoming DDL bursts. In our evaluation, PCN improves the throughput of DDL traffic by 72% on average.
    번역하기

    This thesis presents “Proactive Congestion Notification” (PCN), a congestionavoidance technique for distributed deep learning (DDL). DDL is widely used to scale out and accelerate deep neural network training. In DDL, each worker trains a copy of ...

    This thesis presents “Proactive Congestion Notification” (PCN), a congestionavoidance technique for distributed deep learning (DDL). DDL is widely used to scale out and accelerate deep neural network training. In DDL, each worker trains a copy of the deep learning model with different training inputs and synchronizes the model gradients at the end of each iteration. However, it is well known that the network communication for synchronizing model parameters is the main bottleneck in DDL. Our key observation is that the DDL architecture makes each worker generate burst traffic every iteration, which causes network congestion and in turn degrades the throughput of DDL traffic. Based on this observation, the key idea behind PCN is to prevent potential congestion by proactively regulating the switch queue length before DDL burst traffic arrives at the switch, which prepares the switches for handling incoming DDL bursts. In our evaluation, PCN improves the throughput of DDL traffic by 72% on average.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 2 Background and Motivation 6
    • 2.1 Distributed Deep Learning Traffic 6
    • 2.2 Related Work 10
    • 2.2.1 Host-side scheduling on DDL traffic 10
    • 1 Introduction 1
    • 2 Background and Motivation 6
    • 2.1 Distributed Deep Learning Traffic 6
    • 2.2 Related Work 10
    • 2.2.1 Host-side scheduling on DDL traffic 10
    • 2.2.2 Reducing the amount of DDL Traffic 12
    • 2.2.3 Improvement using an in-network switch 13
    • 2.2.4 ECN-based approach 15
    • 2.2.5 Novelty of PCN 16
    • 2.3 P4 and Switch Programmability 17
    • 3 Design 19
    • 3.1 Communication Sequence between Worker and PS 19
    • 3.2 Switch Operation for PCN 21
    • 3.3 PCN Threshold Policy 24
    • 4 Evaluation 26
    • 4.1 Evaluation Metrics 27
    • 4.2 Queue Length Change 28
    • 4.3 Performance Improvement 30
    • 4.4 Overheads 31
    • 5 Discussion 32
    • 5.1 Estimation of the total DDL training time 32
    • 5.2 The impact of environment changes on PCN 33
    • 5.3 Multi-tenancy 34
    • 6 Conclusion 35
    • Bibliography 36
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼