Resource Aware GPU Scheduling in Kubernetes Infrastructure

descriptionPublicationkeyboard_double_arrow_right Conference object , Article 01 Jan 2021Embargo end date: 02 Mar 2021 English Publisher:Schloss Dagstuhl – Leibniz-Zentrum für InformatikFunded by:EC | EVOLVE

Authors: Ferikoglou, Aggelos; Masouros, Dimosthenis; Tzenetopoulos, Achilleas; Xydis, Sotirios; Soudris, Dimitrios;

Resource Aware GPU Scheduling in Kubernetes Infrastructure

- Summary
- Subjects
- Metrics

Abstract

Nowadays, there is an ever-increasing number of artificial intelligence inference workloads pushed and executed on the cloud. To effectively serve and manage the computational demands, data center operators have provisioned their infrastructures with accelerators. Specifically for GPUs, support for efficient management lacks, as state-of-the-art schedulers and orchestrators, threat GPUs only as typical compute resources ignoring their unique characteristics and application properties. This phenomenon combined with the GPU over-provisioning problem leads to severe resource under-utilization. Even though prior work has addressed this problem by colocating applications into a single accelerator device, its resource agnostic nature does not manage to face the resource under-utilization and quality of service violations especially for latency critical applications. In this paper, we design a resource aware GPU scheduling framework, able to efficiently colocate applications on the same GPU accelerator card. We integrate our solution with Kubernetes, one of the most widely used cloud orchestration frameworks. We show that our scheduler can achieve 58.8% lower end-to-end job execution time 99%-ile, while delivering 52.5% higher GPU memory usage, 105.9% higher GPU utilization percentage on average and 44.4% lower energy consumption on average, compared to the state-of-the-art schedulers, for a variety of ML representative workloads.

Related Organizations

Harokopio University
Greece
Leibniz Association
Germany
National Technical University of Athens
Greece
Schloss Dagstuhl – Leibniz Center for Informatics
Germany

Keywords

Computer systems organization → Heterogeneous (hybrid) systems, GPU scheduling, Hardware → Emerging architectures, cloud computing, Computer systems organization → Cloud computing, heterogeneity, kubernetes, Computing methodologies, 004, ddc: ddc:004

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

0

Average

Green

Funded by

EC| EVOLVE