RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Research on GPU scheduling and resource sharing for GPU-centric computing model

    한글로보기

    https://www.riss.kr/link?id=T16974301

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent advancements in parallel computing have positioned the Graphics Processing Unit (GPU) as a key computing unit in multi-tenant and multi-application environments. Despite the shift towards GPU-centric computing, current GPU resource sharing technologies face limitations due to the inability to facilitate communication between divided resources, thus restricting efficient resource sharing and interaction. Additionally, the inability of operating systems to control GPU tasks exacerbates priority inversion problems, preventing high-priority kernels from preempting lower-priority tasks.
    Addressing these challenges, this paper proposes a GPU scheduling and resource sharing model, controlled by the operating system, to enhance GPU-centric computing. Our model introduces three key schemes: GPU Kernel Transactionization, GPU Kernel Idempotent Classification, and Distributed Training with MIG.
    The GPU Kernel Transactionization scheme defines the execution transaction of the kernel for GPU preemptive scheduling. This scheme ensures data consistency of the kernel buffer, which changes during kernel execution, through GPU memory snapshot/rollback technology and a definition of a single GPU kernel execution transaction. Through this scheme, a preemption delay of 18μs within the 99.9th percentile is guaranteed, and in most cases, the average execution time delay of high-priority tasks is less than 10% even in an overloaded state.
    The GPU Kernel Idempotent Classify scheme utilizes static analysis to detect idempotent GPU kernels, which do not damage input data during execution, and employs them in preemptive scheduling. The detected idempotent kernel information is dynamically conveyed to the GPU scheduler to perform preemptive scheduling without memory context snapshot/rollback costs. Through this scheme, 12 out of 22 kernels in the Rodinia benchmark suite were detected as idempotent kernels, and preemptive scheduling could be applied to 63.6% of applications without memory context switching costs.
    The Distributed Training with MIG scheme supports communication between MIG instances for distributed training on both multiple and single GPUs. This method improves GPU resource efficiency by allowing concurrent tasks and increases computational parallelism within a single GPU. Using this scheme, an average increase of 14% in throughput was achieved when conducting distributed training of multiple models simultaneously, and a maximum performance improvement of 21.9% was realized when transitioning from single to distributed training.
    Through these three schemes, we provide preemptive scheduling support for GPU tasks in the operating system and an efficient GPU resource sharing method for DNN training. Through this, we construct a GPU resource sharing computing system that uses the GPU as a primary computing unit.
    번역하기

    Recent advancements in parallel computing have positioned the Graphics Processing Unit (GPU) as a key computing unit in multi-tenant and multi-application environments. Despite the shift towards GPU-centric computing, current GPU resource sharing tech...

    Recent advancements in parallel computing have positioned the Graphics Processing Unit (GPU) as a key computing unit in multi-tenant and multi-application environments. Despite the shift towards GPU-centric computing, current GPU resource sharing technologies face limitations due to the inability to facilitate communication between divided resources, thus restricting efficient resource sharing and interaction. Additionally, the inability of operating systems to control GPU tasks exacerbates priority inversion problems, preventing high-priority kernels from preempting lower-priority tasks.
    Addressing these challenges, this paper proposes a GPU scheduling and resource sharing model, controlled by the operating system, to enhance GPU-centric computing. Our model introduces three key schemes: GPU Kernel Transactionization, GPU Kernel Idempotent Classification, and Distributed Training with MIG.
    The GPU Kernel Transactionization scheme defines the execution transaction of the kernel for GPU preemptive scheduling. This scheme ensures data consistency of the kernel buffer, which changes during kernel execution, through GPU memory snapshot/rollback technology and a definition of a single GPU kernel execution transaction. Through this scheme, a preemption delay of 18μs within the 99.9th percentile is guaranteed, and in most cases, the average execution time delay of high-priority tasks is less than 10% even in an overloaded state.
    The GPU Kernel Idempotent Classify scheme utilizes static analysis to detect idempotent GPU kernels, which do not damage input data during execution, and employs them in preemptive scheduling. The detected idempotent kernel information is dynamically conveyed to the GPU scheduler to perform preemptive scheduling without memory context snapshot/rollback costs. Through this scheme, 12 out of 22 kernels in the Rodinia benchmark suite were detected as idempotent kernels, and preemptive scheduling could be applied to 63.6% of applications without memory context switching costs.
    The Distributed Training with MIG scheme supports communication between MIG instances for distributed training on both multiple and single GPUs. This method improves GPU resource efficiency by allowing concurrent tasks and increases computational parallelism within a single GPU. Using this scheme, an average increase of 14% in throughput was achieved when conducting distributed training of multiple models simultaneously, and a maximum performance improvement of 21.9% was realized when transitioning from single to distributed training.
    Through these three schemes, we provide preemptive scheduling support for GPU tasks in the operating system and an efficient GPU resource sharing method for DNN training. Through this, we construct a GPU resource sharing computing system that uses the GPU as a primary computing unit.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 병렬 컴퓨팅의 발전으로, 그래픽 처리 장치(GPU)가 멀티-테넌트 및 멀티-애플리케이션 환경에서 주요 컴퓨팅 단위로 자리 잡고 있다. GPU 중심의 컴퓨팅으로의 전환에도 불구하고, 현재의 GPU 자원 공유 기술은 분할된 자원 간 통신이 불가능하여, 효율적인 자원 공유와 상호 작용에 제한을 두고 있다. 또한, 운영체제가 GPU 작업을 제어하지 못하는 문제는 우선순위 역전 문제를 악화시키며, 높은 우선순위의 커널이 낮은 우선순위의 작업을 선점하지 못한다.
    이러한 문제를 해결하기 위해, 본 논문에서는 운영체제가 제어하는 새로운 GPU 스케줄링 및 자원 공유 모델을 제안하여 GPU 중심의 컴퓨팅을 강화한다. 우리의 모델은 GPU 커널 트랜잭션화, GPU 커널 멱등성 분류, 그리고 MIG 기반 분산 학습이라는 세 가지 주요 기법을 도입한다.
    먼저 GPU Kernel Transactionization 기법은 GPU 선점형 스케줄링을 위한 커널의 실행 트랜잭션을 정의한다. 이 기법은 GPU 메모리 스냅샷/롤백 기술과 하나의 GPU 커널 실행 트랜잭션 정의를 통해 커널 실행 시 변경되는 커널 버퍼의 데이터 일관성을 보장한다. 본 기법을 통해 99.9th percentail 내에서 18μs의 선점 지연을 보장하며, 대부분의 경우 과부하 상태에서도 우선순위가 높은 작업의 실행 시간 평균 지연은 10% 미만이다.
    GPU Kernel Idempotent classify 기법은 정적 분석을 통해 실행 중 입력 데이터를 손상시키지 않는 명등성 GPU 커널을 검출하여 선점형 스케줄링에 활용한다. 검출된 멱등성 커널 정보는 동적으로 GPU 스케줄러에 전달하여 메모리 컨텍스트 스냅샷/롤백 비용 없이 선점형 스케줄링을 수행한다. 본 기법을 통해 rodinia benchmark suite에서 22의 커널 중 12개의 커널이 멱등성 커널로 검출되었으며, 63.6% 비중의 애플리케이션에 대해 메모리 컨텍스트 스위칭 비용 없이 선점형 스케줄링을 적용할 수 있었다.
    마지막으로 Distributed Training with MIG 기법은 여러 GPU 뿐만 아니라 단일 GPU에서도 MIG 인스턴스 간의 통신을 지원한다. 이 방법은 여러 작업이 동시에 GPU 자원을 사용할 수 있게 하여 GPU 자원의 효율성을 향상시킨다. 또한 단일 GPU 내의 인스턴스를 사용하여 분산 학습을 함으로써 계산 병렬성을 증가시켜 학습 성능을 개선한다. 이 방식을 통해 여러 모델을 동시에 분산 학습할 때 처리량이 평균 14% 증가하고, 단일 학습에서 분산 학습으로 전환할 때 최대 21.9%의 성능 향상을 달성할 수 있었다.
    우리는 세 가지의 기법을 통해서 운영체제에서 GPU 작업에 대한 선점형 스케줄링 지원과 DNN 프레임워크에서 효율적인 GPU 자원 공유 방법을 제공한다. 우리는 이것을 통해 GPU를 primary computing unit으로 사용하는 GPU 자원 공유 컴퓨팅 시스템을 구축한다.
    번역하기

    최근 병렬 컴퓨팅의 발전으로, 그래픽 처리 장치(GPU)가 멀티-테넌트 및 멀티-애플리케이션 환경에서 주요 컴퓨팅 단위로 자리 잡고 있다. GPU 중심의 컴퓨팅으로의 전환에도 불구하고, 현재의...

    최근 병렬 컴퓨팅의 발전으로, 그래픽 처리 장치(GPU)가 멀티-테넌트 및 멀티-애플리케이션 환경에서 주요 컴퓨팅 단위로 자리 잡고 있다. GPU 중심의 컴퓨팅으로의 전환에도 불구하고, 현재의 GPU 자원 공유 기술은 분할된 자원 간 통신이 불가능하여, 효율적인 자원 공유와 상호 작용에 제한을 두고 있다. 또한, 운영체제가 GPU 작업을 제어하지 못하는 문제는 우선순위 역전 문제를 악화시키며, 높은 우선순위의 커널이 낮은 우선순위의 작업을 선점하지 못한다.
    이러한 문제를 해결하기 위해, 본 논문에서는 운영체제가 제어하는 새로운 GPU 스케줄링 및 자원 공유 모델을 제안하여 GPU 중심의 컴퓨팅을 강화한다. 우리의 모델은 GPU 커널 트랜잭션화, GPU 커널 멱등성 분류, 그리고 MIG 기반 분산 학습이라는 세 가지 주요 기법을 도입한다.
    먼저 GPU Kernel Transactionization 기법은 GPU 선점형 스케줄링을 위한 커널의 실행 트랜잭션을 정의한다. 이 기법은 GPU 메모리 스냅샷/롤백 기술과 하나의 GPU 커널 실행 트랜잭션 정의를 통해 커널 실행 시 변경되는 커널 버퍼의 데이터 일관성을 보장한다. 본 기법을 통해 99.9th percentail 내에서 18μs의 선점 지연을 보장하며, 대부분의 경우 과부하 상태에서도 우선순위가 높은 작업의 실행 시간 평균 지연은 10% 미만이다.
    GPU Kernel Idempotent classify 기법은 정적 분석을 통해 실행 중 입력 데이터를 손상시키지 않는 명등성 GPU 커널을 검출하여 선점형 스케줄링에 활용한다. 검출된 멱등성 커널 정보는 동적으로 GPU 스케줄러에 전달하여 메모리 컨텍스트 스냅샷/롤백 비용 없이 선점형 스케줄링을 수행한다. 본 기법을 통해 rodinia benchmark suite에서 22의 커널 중 12개의 커널이 멱등성 커널로 검출되었으며, 63.6% 비중의 애플리케이션에 대해 메모리 컨텍스트 스위칭 비용 없이 선점형 스케줄링을 적용할 수 있었다.
    마지막으로 Distributed Training with MIG 기법은 여러 GPU 뿐만 아니라 단일 GPU에서도 MIG 인스턴스 간의 통신을 지원한다. 이 방법은 여러 작업이 동시에 GPU 자원을 사용할 수 있게 하여 GPU 자원의 효율성을 향상시킨다. 또한 단일 GPU 내의 인스턴스를 사용하여 분산 학습을 함으로써 계산 병렬성을 증가시켜 학습 성능을 개선한다. 이 방식을 통해 여러 모델을 동시에 분산 학습할 때 처리량이 평균 14% 증가하고, 단일 학습에서 분산 학습으로 전환할 때 최대 21.9%의 성능 향상을 달성할 수 있었다.
    우리는 세 가지의 기법을 통해서 운영체제에서 GPU 작업에 대한 선점형 스케줄링 지원과 DNN 프레임워크에서 효율적인 GPU 자원 공유 방법을 제공한다. 우리는 이것을 통해 GPU를 primary computing unit으로 사용하는 GPU 자원 공유 컴퓨팅 시스템을 구축한다.

    더보기

    목차 (Table of Contents)

    • Chapter 1. Introduction 1
    • 1.1 GPU Computing System 1
    • 1.2 Preemptive Scheduling for GPU 4
    • 1.3 GPU Resource Sharing for Multi-Tasking 7
    • 1.4 Organization 10
    • Chapter 1. Introduction 1
    • 1.1 GPU Computing System 1
    • 1.2 Preemptive Scheduling for GPU 4
    • 1.3 GPU Resource Sharing for Multi-Tasking 7
    • 1.4 Organization 10
    • Chapter 2. Background and Related Work 11
    • 2.1 OpenCL Runtime Model on HSA 11
    • 2.2 GPU Kernel Idempotent 13
    • 2.3 GPU Preemptive Scheduling 16
    • 2.4 GPU Resource Sharing for Deep Learning 18
    • Chapter 3. A GPU Kernel Transactionization Scheme 25
    • 3.1 Overview 25
    • 3.2 Kernel Snapshotting for Transactionization 28
    • 3.3 Transactionization Process 31
    • 3.4 Evaluation 34
    • 3.5 Summary 48
    • Chapter 4. Idempotent-Based Preemptive GPU Kernel Scheduling 49
    • 4.1 Overview 49
    • 4.2 Idempotent Kernel Classification 51
    • 4.3 A GPU Kernel Transactionization with Idempotent 57
    • 4.4 Priority-based GPU Preemptive Scheduling 58
    • 4.5 Evaluation 63
    • 4.6 Summary 80
    • Chapter 5. Distributed Training with MIG 81
    • 5.1 Overview 81
    • 5.2 MIG Context Manager 83
    • 5.3 Communication for Distributed Training with MIG 86
    • 5.4 Evaluation 90
    • 5.5 Summary 105
    • Chapter 6. Conclusion 107
    • References 109
    • Korean Abstract 119
    더보기

    참고문헌 (Reference)

    1. 9 OpenCL, Khronos Group, Online Available https www. khronos. org/opencl, , 2009

    2. NVIDIA Tesla P100, NVIDIA, White Paper, , 2016

    3. Exynos specification, Samsung Electronics Co., Ltd., Online Available https//www. samsung. com/semiconductor/minisite/exynos/, , 2019

    4. MULTIPROCESS SERVICE, NVIDIA, White paper, , 2022

    5. NVIDIA GeForce GTX 1080, NVIDIA, White Paper, , 2016

    6. NVIDIA GPUDirect Storage, NVIDIA, White Paper, , 2021

    7. ThroughSilicon Via (TSV), M. Motoyoshi, Proc. IEEE, vol. 97, no. 1, pp. 43–48, , 2009

    8. Operating System Concepts, Baer and, G. Greg, G. Peter, S. Abraham, 9th ed Hoboken, NJ, USA: Wiley, Art. no. 271, , 2013

    9. The Definitive ANTLR 4 Reference, T. Parr, Pragmatic Bookshelf, , 2013

    10. 8 Accelerated Linear Algebra(XLA), Google Tensorflow, Online Available https//www tensorflow org/xla, , 2017

    1. 9 OpenCL, Khronos Group, Online Available https www. khronos. org/opencl, , 2009

    2. NVIDIA Tesla P100, NVIDIA, White Paper, , 2016

    3. Exynos specification, Samsung Electronics Co., Ltd., Online Available https//www. samsung. com/semiconductor/minisite/exynos/, , 2019

    4. MULTIPROCESS SERVICE, NVIDIA, White paper, , 2022

    5. NVIDIA GeForce GTX 1080, NVIDIA, White Paper, , 2016

    6. NVIDIA GPUDirect Storage, NVIDIA, White Paper, , 2021

    7. ThroughSilicon Via (TSV), M. Motoyoshi, Proc. IEEE, vol. 97, no. 1, pp. 43–48, , 2009

    8. Operating System Concepts, Baer and, G. Greg, G. Peter, S. Abraham, 9th ed Hoboken, NJ, USA: Wiley, Art. no. 271, , 2013

    9. The Definitive ANTLR 4 Reference, T. Parr, Pragmatic Bookshelf, , 2013

    10. 8 Accelerated Linear Algebra(XLA), Google Tensorflow, Online Available https//www tensorflow org/xla, , 2017

    11. Heterogeneous system architecture, Heterogeneous System Architecture Foundation, Online Available http//hsafoundation com, , 2016

    12. Idempotent processor architecture, M. De Kruijf and, K. Sankaralingam, in Proc 44th Annu. IEEE/ACM Int. Symp. Microarchit., pp. 140–151, , 2011

    13. 57 Getting started with CUDA graphs, NVIDIA, Online Available https://developer. nvidia. com/blog/cudagraphs/, , 2019

    14. Measuring function duration with ftrace, T. Bird, in Proceedings of the Linux Symposium. Citeseer, pp. 47–54, , 2009

    15. NVIDIA A100 Tensor Core GPU Architecture, NVIDIA, White paper, , 2020

    16. NVIDIA MultiInstance GPU user guide 2023, NVIDIA, Online Available https://docs. nvidia. com/datacenter/tesla/miguserguide/index. html, , 2023

    17. Exploring memory persistency models for GPUs, H. Zhou, Y. Solihin and, Z. Lin, M. Alshboul, in Proc. IEEE 28th Int. Conf. Parallel Archit. Compilation Techn., pp. 311–323, , 2019

    18. Enabling preemptive multiprogramming on GPUs,, I. Gelado, N. Navarro and, M. Valero, J. Cabezas, I. Tanasic, A. Ramirez, in Proceedings of 41st Annu. ACM/IEEE Int. Symp. Comput. Archit., pp. 193–204, , 2014

    19. Improving GPGPU concurrency with elastic kernels, Ramaswamy Govindarajan, Matthew J. Thazhuthaveetil and, Pai, Sreepathi, 41.1, 407418, , 2013

    20. Multitasking realtime embedded GPU computing tasks, Pιnar MuyanÖzçelik and, J. D. Owens, in Proceedings of 7th Int. Workshop Program. Models Appl. Multicores Manycores, pp. 78–87, , 2016

    21. Benchmarking and analyzing deep neural network training, Zhu, Hongyu, et al, in the Proceedings of 2018 IEEE International Symposium on Workload Characterization (IISWC 2018, , 2018

    22. FLEP: Enabling flexible and efficient preemption on GPUs, B. Wu, X. Liu, X. Zhou and, C. Jiang, in Proceedings of 22nd Int. Conf. Archit. Support Program. Lang. Operating Syst., pp. 483–496, , 2017

    23. Operating systems challenges for GPU resource management, S. Kato, R. Rajkumar, S. Brandt, Y. Ishikawa and, in Proc. Int. Workshop Operating Syst. Platforms Embedded RealTime Appl., pp. 23–32, , 2011

    24. iGPU: Exception support and speculative execution on GPUs, M. De Kruijf and, K. Sankaralingam, J. Menon, in Proceedings of 39th Annu. Int. Symp. Comput. Archit., pp. 72–83, , 2018

    25. Checkpointing and rollbackrecovery for distributed systems, R. Koo and, S. Toueg, IEEE Trans. Softw. Eng., vol. SE13, no. 1, pp. 23–31, , 1987

    26. Open source mali midgard GPUkernel drivers (rel. r5p006rel0, ARMCo., Ltd., Online Available https://developer. arm. com/toolsandsoftware/graphicsandgaming/malidrivers/midgardkernel, , 2014

    27. An analysis of collocation on GPUs for deep learning training, Ehsan YousefzadehAslMiandoab and, Pınar Tözün, Robroek, Ties, arXiv eprints arXiv2209, , 2022

    28. Static analysis and compiler design for idempotent processing, S. Jha, M. De Kruijf, K. Sankaralingam and, in Proceedings of 33rd ACM SIGPLANConf. Program. Lang. Des. Implementation, pp. 475–486, , 2012

    29. iDO: compilerdirected failure atomicity for nonvolatile memory, S. H. Noh and, S. K. Lee, Q. Liu, M. L. Scott, C. Jung, J. Izraelevitz, in Proceedings of 51st Annu. IEEE/ACM Int. Symp. Microarchit., pp. 258–270, , 2018

    30. Characterizing multiinstance GPU for machine learning workloads, Li, Baolin, et al, in the Proceedings of 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW 2022, , 2022

    31. Supporting preemptive task executions and memory copies in GPGPUs, C. Basaran and, K. Kang, in Proc 24th Euromicro Conf. RealTime Syst., pp. 287–296, , 2012

    32. Chimera: Collaborative preemption for multitasking on a shared GPU, S. Mahlke, J. Park, Y. Park and, J. Kyu, in Proc. 20th Int. Conf. Archit. Support Program. Lang. Operating Syst., pp. 593–606, , 2015

    33. Idempotent code generation: Implementation, analysis, and evaluation, K. Sankaralingam, M. De Kruijf and, in Proceedings of IEEE/ACM Int. Symp. Code Gener. Optim., pp. 1–12, , 2013

    34. 41 NVIDIA’s next generation CUDA compute architecture: kepler GK110, NVIDIA, White Paper, , 2012

    35. Accelerating adaptive background modeling on lowpower integrated GPUs, S. Wills, S. Azmat, L. Wills and, in Proceedings of 41st Int. Conf. Parallel Process. Workshops, pp. 568–573, , 2012

    36. Nimble: Lightweight and parallel gpu task scheduling for deep learning, Kwon, Woosuk, et al, Advances in Neural Information Processing Systems, 33, 83438354, , 2020

    37. PTask: Operating system abstractions to manage GPUs as compute devices, B. Ray and, M. Silberstein, C. J. Rossbach, J. Currey, E. Witchel, in Proceedings of 23rd ACM Symp. Operating Syst. Princ. pp. 233–248, , 2011

    38. Page placement strategies for GPUs within heterogeneous memory systems, S. W. Keckler, N. Agarwal, M. O’Connor and, D. Nellans, M. Stephenson, in Proceedings of 20th Int. Conf. Archit. Support Program. Lang. Operating Syst., , 2015

    39. Towards adaptive GPU resource management for embedded realtime systems, S. Kato, R. R. Rajkumar and, J. Kim, ACM SIGBED Rev., vol. 10, no. 1, pp. 14–17, , 2013

    40. General purpose computing on lowpower embedded GPUs: Has it come of age?, Z. Peng, U. D. Bordoloi, A. Maghazeh, P. Eles and, in Proceedings of Int. Conf. Embedded Comput. Syst., Archite. Model. Simul., pp. 1–10, , 2013

    41. Repurposing GPU microarchitectures with lightweight outoforder execution, K. Iliakis, S. Xydis and, D. Soudris, in IEEE Transactions on Parallel and Distributed Systems, , 2022

    42. Salus: Finegrained GPU sharing primitives for deep learning applications, Mosharaf Chowdhury, Yu, Peifeng and, arXiv preprint arXiv:1902.04610, , 2019

    43. Efficient checkpointing with recompute scheme for nonvolatile main memory, K. Kimura, R. Elkhouly, H. Elnawawy, Y. Solihin, J. Tuck and, M. Alshboul, ACM Trans. Archit. Code Optim., vol. 16, no. 2, , 2019

    44. An efficient checkpoint and recovery mechanism for realtime embedded systems, Y. Bai, Q. Chen, C. Wang, J. Zeng and, G. Luan, in Proceedings of IEEE Int. Conf. Parallel Distrib. Process. Appl. Ubiquitous Comput. Commun. Big Data Cloud Comput. Social Comput. Netw. Sustain. Comput. Commun., pp. 824 831, , 2018

    45. Salus: Finegrained gpu sharing primitives for deep learning applications, in, P. Yu and, M. Chowdhury, Proceedings of Machine Learning and Systems 2020 (MLSys), , 2020

    46. EffiSha: A software framework for enabling effficient preemptive scheduling of GPU, X. Shen and, Y. Zhao, H. Zhou, G. Chen, in Proceedings of 22nd ACM SIGPLAN Symp. Princ. Practice Parallel Program., pp. 3–16, , 2017

    47. Kernelet: Highthroughput GPU kernel executions with dynamic slicing and scheduling, Zhong, Jianlong and, Bingsheng He, IEEE Transactions on Parallel and Distributed Systems 25.6, 15221532, , 2013

    48. HPC driven innovations in network management for greater efficiency and productivity, Veerla, Harsha, et al, in the Proceedings of 5th International Conference on Inventive Research in Computing Applications (ICIRCA 2023, , 2023

    49. Enabling efficient preemption for SIMT architectures with lightweight context switching, L. Nyland and, H. Zhou, Z. Lin, in Proceedings of Int. Conf. High Perform. Comput. Netw. Storage Anal., pp. 898–908, , 2016

    50. MuxFlow: Efficient and safe GPU sharing in largescale production deep learning clusters, Zhao, Yihao, et al, arXiv preprint arXiv:2303.13803 2023, , 2023

    51. Compilerdirected lightweight checkpointing for finegrained guaranteed soft error recovery, D. Tiwari, Q. Liu, D. Lee and, C. Jung, in Proceedings of Int. Conf. High Perform. Comput. Netw. Storage Anal., pp. 228–239, , 2016

    52. Serving heterogeneous machine learning models on multiGPU servers with spatiotemporal sharing, Seungbeom Choi, et al, in the Proceedings of 2022 USENIX Annual Technical Conference (ATC 22), , 2022

    53. Kubeknots: Resource harvesting through dynamic container orchestration in GPUbased datacenters, P. Thinakaran, J. R. Gunasekaran, et al., in Proceedings of 2019 IEEE International Conference on Cluster Computing (CLUSTER 19, , 2019

    54. Towards multitenant GPGPU: Eventdriven programming model for systemwide scheduling on shared GPUs, H. Yamada, S. Kato and, K. Kono, Y. Suzuki, in Proceedings of Workshop Multicore RackScale Syst., , 2016

    55. Serving DNN models with multiinstance gpus: A case of the reconfigurable machine scheduling problem, Tan, Cheng, et al, arXiv preprint arXiv:2109.11067, , 2021

    56. Warpedslicer: efficient intrasm slicing through dynamic resource partitioning for gpu multiprogramming, Xu, Qiumin, et al, , 230242, , 2016

    57. Research on electronic hardware scheme design for performance improvement of convolutional neural network, Liu, Xibin, Xiaofang Liu, Han Zhu and, 2023 IEEE 3rd International Conference on Electronic Technology, Communication and Information (ICETCI 23, , 2023

    58. The design and implementation of berkeley lab’s Linux checkpoint/restart16 Adaptive dynamic checkpointing for safe efficient intermittent computing, P. Hargrove and, E. Roman, B. Lucia, K. Maeng and, J. Duell, in Proceedings of 13th USENIX Symp. Operating Syst. Des. Implementation, pp. 129 144, , 2002

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼