1. 9 OpenCL, Khronos Group, Online Available https www. khronos. org/opencl, , 2009
2. NVIDIA Tesla P100, NVIDIA, White Paper, , 2016
3. Exynos specification, Samsung Electronics Co., Ltd., Online Available https//www. samsung. com/semiconductor/minisite/exynos/, , 2019
4. MULTIPROCESS SERVICE, NVIDIA, White paper, , 2022
5. NVIDIA GeForce GTX 1080, NVIDIA, White Paper, , 2016
6. NVIDIA GPUDirect Storage, NVIDIA, White Paper, , 2021
7. ThroughSilicon Via (TSV), M. Motoyoshi, Proc. IEEE, vol. 97, no. 1, pp. 43–48, , 2009
8. Operating System Concepts, Baer and, G. Greg, G. Peter, S. Abraham, 9th ed Hoboken, NJ, USA: Wiley, Art. no. 271, , 2013
9. The Definitive ANTLR 4 Reference, T. Parr, Pragmatic Bookshelf, , 2013
10. 8 Accelerated Linear Algebra(XLA), Google Tensorflow, Online Available https//www tensorflow org/xla, , 2017
11. Heterogeneous system architecture, Heterogeneous System Architecture Foundation, Online Available http//hsafoundation com, , 2016
12. Idempotent processor architecture, M. De Kruijf and, K. Sankaralingam, in Proc 44th Annu. IEEE/ACM Int. Symp. Microarchit., pp. 140–151, , 2011
13. 57 Getting started with CUDA graphs, NVIDIA, Online Available https://developer. nvidia. com/blog/cudagraphs/, , 2019
14. Measuring function duration with ftrace, T. Bird, in Proceedings of the Linux Symposium. Citeseer, pp. 47–54, , 2009
15. NVIDIA A100 Tensor Core GPU Architecture, NVIDIA, White paper, , 2020
16. NVIDIA MultiInstance GPU user guide 2023, NVIDIA, Online Available https://docs. nvidia. com/datacenter/tesla/miguserguide/index. html, , 2023
17. Exploring memory persistency models for GPUs, H. Zhou, Y. Solihin and, Z. Lin, M. Alshboul, in Proc. IEEE 28th Int. Conf. Parallel Archit. Compilation Techn., pp. 311–323, , 2019
18. Enabling preemptive multiprogramming on GPUs,, I. Gelado, N. Navarro and, M. Valero, J. Cabezas, I. Tanasic, A. Ramirez, in Proceedings of 41st Annu. ACM/IEEE Int. Symp. Comput. Archit., pp. 193–204, , 2014
19. Improving GPGPU concurrency with elastic kernels, Ramaswamy Govindarajan, Matthew J. Thazhuthaveetil and, Pai, Sreepathi, 41.1, 407418, , 2013
20. Multitasking realtime embedded GPU computing tasks, Pιnar MuyanÖzçelik and, J. D. Owens, in Proceedings of 7th Int. Workshop Program. Models Appl. Multicores Manycores, pp. 78–87, , 2016
21. Benchmarking and analyzing deep neural network training, Zhu, Hongyu, et al, in the Proceedings of 2018 IEEE International Symposium on Workload Characterization (IISWC 2018, , 2018
22. FLEP: Enabling flexible and efficient preemption on GPUs, B. Wu, X. Liu, X. Zhou and, C. Jiang, in Proceedings of 22nd Int. Conf. Archit. Support Program. Lang. Operating Syst., pp. 483–496, , 2017
23. Operating systems challenges for GPU resource management, S. Kato, R. Rajkumar, S. Brandt, Y. Ishikawa and, in Proc. Int. Workshop Operating Syst. Platforms Embedded RealTime Appl., pp. 23–32, , 2011
24. iGPU: Exception support and speculative execution on GPUs, M. De Kruijf and, K. Sankaralingam, J. Menon, in Proceedings of 39th Annu. Int. Symp. Comput. Archit., pp. 72–83, , 2018
25. Checkpointing and rollbackrecovery for distributed systems, R. Koo and, S. Toueg, IEEE Trans. Softw. Eng., vol. SE13, no. 1, pp. 23–31, , 1987
26. Open source mali midgard GPUkernel drivers (rel. r5p006rel0, ARMCo., Ltd., Online Available https://developer. arm. com/toolsandsoftware/graphicsandgaming/malidrivers/midgardkernel, , 2014
27. An analysis of collocation on GPUs for deep learning training, Ehsan YousefzadehAslMiandoab and, Pınar Tözün, Robroek, Ties, arXiv eprints arXiv2209, , 2022
28. Static analysis and compiler design for idempotent processing, S. Jha, M. De Kruijf, K. Sankaralingam and, in Proceedings of 33rd ACM SIGPLANConf. Program. Lang. Des. Implementation, pp. 475–486, , 2012
29. iDO: compilerdirected failure atomicity for nonvolatile memory, S. H. Noh and, S. K. Lee, Q. Liu, M. L. Scott, C. Jung, J. Izraelevitz, in Proceedings of 51st Annu. IEEE/ACM Int. Symp. Microarchit., pp. 258–270, , 2018
30. Characterizing multiinstance GPU for machine learning workloads, Li, Baolin, et al, in the Proceedings of 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW 2022, , 2022
31. Supporting preemptive task executions and memory copies in GPGPUs, C. Basaran and, K. Kang, in Proc 24th Euromicro Conf. RealTime Syst., pp. 287–296, , 2012
32. Chimera: Collaborative preemption for multitasking on a shared GPU, S. Mahlke, J. Park, Y. Park and, J. Kyu, in Proc. 20th Int. Conf. Archit. Support Program. Lang. Operating Syst., pp. 593–606, , 2015
33. Idempotent code generation: Implementation, analysis, and evaluation, K. Sankaralingam, M. De Kruijf and, in Proceedings of IEEE/ACM Int. Symp. Code Gener. Optim., pp. 1–12, , 2013
34. 41 NVIDIA’s next generation CUDA compute architecture: kepler GK110, NVIDIA, White Paper, , 2012
35. Accelerating adaptive background modeling on lowpower integrated GPUs, S. Wills, S. Azmat, L. Wills and, in Proceedings of 41st Int. Conf. Parallel Process. Workshops, pp. 568–573, , 2012
36. Nimble: Lightweight and parallel gpu task scheduling for deep learning, Kwon, Woosuk, et al, Advances in Neural Information Processing Systems, 33, 83438354, , 2020
37. PTask: Operating system abstractions to manage GPUs as compute devices, B. Ray and, M. Silberstein, C. J. Rossbach, J. Currey, E. Witchel, in Proceedings of 23rd ACM Symp. Operating Syst. Princ. pp. 233–248, , 2011
38. Page placement strategies for GPUs within heterogeneous memory systems, S. W. Keckler, N. Agarwal, M. O’Connor and, D. Nellans, M. Stephenson, in Proceedings of 20th Int. Conf. Archit. Support Program. Lang. Operating Syst., , 2015
39. Towards adaptive GPU resource management for embedded realtime systems, S. Kato, R. R. Rajkumar and, J. Kim, ACM SIGBED Rev., vol. 10, no. 1, pp. 14–17, , 2013
40. General purpose computing on lowpower embedded GPUs: Has it come of age?, Z. Peng, U. D. Bordoloi, A. Maghazeh, P. Eles and, in Proceedings of Int. Conf. Embedded Comput. Syst., Archite. Model. Simul., pp. 1–10, , 2013
41. Repurposing GPU microarchitectures with lightweight outoforder execution, K. Iliakis, S. Xydis and, D. Soudris, in IEEE Transactions on Parallel and Distributed Systems, , 2022
42. Salus: Finegrained GPU sharing primitives for deep learning applications, Mosharaf Chowdhury, Yu, Peifeng and, arXiv preprint arXiv:1902.04610, , 2019
43. Efficient checkpointing with recompute scheme for nonvolatile main memory, K. Kimura, R. Elkhouly, H. Elnawawy, Y. Solihin, J. Tuck and, M. Alshboul, ACM Trans. Archit. Code Optim., vol. 16, no. 2, , 2019
44. An efficient checkpoint and recovery mechanism for realtime embedded systems, Y. Bai, Q. Chen, C. Wang, J. Zeng and, G. Luan, in Proceedings of IEEE Int. Conf. Parallel Distrib. Process. Appl. Ubiquitous Comput. Commun. Big Data Cloud Comput. Social Comput. Netw. Sustain. Comput. Commun., pp. 824 831, , 2018
45. Salus: Finegrained gpu sharing primitives for deep learning applications, in, P. Yu and, M. Chowdhury, Proceedings of Machine Learning and Systems 2020 (MLSys), , 2020
46. EffiSha: A software framework for enabling effficient preemptive scheduling of GPU, X. Shen and, Y. Zhao, H. Zhou, G. Chen, in Proceedings of 22nd ACM SIGPLAN Symp. Princ. Practice Parallel Program., pp. 3–16, , 2017
47. Kernelet: Highthroughput GPU kernel executions with dynamic slicing and scheduling, Zhong, Jianlong and, Bingsheng He, IEEE Transactions on Parallel and Distributed Systems 25.6, 15221532, , 2013
48. HPC driven innovations in network management for greater efficiency and productivity, Veerla, Harsha, et al, in the Proceedings of 5th International Conference on Inventive Research in Computing Applications (ICIRCA 2023, , 2023
49. Enabling efficient preemption for SIMT architectures with lightweight context switching, L. Nyland and, H. Zhou, Z. Lin, in Proceedings of Int. Conf. High Perform. Comput. Netw. Storage Anal., pp. 898–908, , 2016
50. MuxFlow: Efficient and safe GPU sharing in largescale production deep learning clusters, Zhao, Yihao, et al, arXiv preprint arXiv:2303.13803 2023, , 2023
51. Compilerdirected lightweight checkpointing for finegrained guaranteed soft error recovery, D. Tiwari, Q. Liu, D. Lee and, C. Jung, in Proceedings of Int. Conf. High Perform. Comput. Netw. Storage Anal., pp. 228–239, , 2016
52. Serving heterogeneous machine learning models on multiGPU servers with spatiotemporal sharing, Seungbeom Choi, et al, in the Proceedings of 2022 USENIX Annual Technical Conference (ATC 22), , 2022
53. Kubeknots: Resource harvesting through dynamic container orchestration in GPUbased datacenters, P. Thinakaran, J. R. Gunasekaran, et al., in Proceedings of 2019 IEEE International Conference on Cluster Computing (CLUSTER 19, , 2019
54. Towards multitenant GPGPU: Eventdriven programming model for systemwide scheduling on shared GPUs, H. Yamada, S. Kato and, K. Kono, Y. Suzuki, in Proceedings of Workshop Multicore RackScale Syst., , 2016
55. Serving DNN models with multiinstance gpus: A case of the reconfigurable machine scheduling problem, Tan, Cheng, et al, arXiv preprint arXiv:2109.11067, , 2021
56. Warpedslicer: efficient intrasm slicing through dynamic resource partitioning for gpu multiprogramming, Xu, Qiumin, et al, , 230242, , 2016
57. Research on electronic hardware scheme design for performance improvement of convolutional neural network, Liu, Xibin, Xiaofang Liu, Han Zhu and, 2023 IEEE 3rd International Conference on Electronic Technology, Communication and Information (ICETCI 23, , 2023
58. The design and implementation of berkeley lab’s Linux checkpoint/restart16 Adaptive dynamic checkpointing for safe efficient intermittent computing, P. Hargrove and, E. Roman, B. Lucia, K. Maeng and, J. Duell, in Proceedings of 13th USENIX Symp. Operating Syst. Des. Implementation, pp. 129 144, , 2002