Downloads provided by UsageCounts
doi: 10.1145/3606376.3593555 , 10.1145/3578338.3593555 , 10.1145/3570609 , 10.48550/arxiv.2210.15315
arXiv: 2210.15315
handle: 11573/1661239 , 11573/1704623 , 20.500.11850/589368
doi: 10.1145/3606376.3593555 , 10.1145/3578338.3593555 , 10.1145/3570609 , 10.48550/arxiv.2210.15315
arXiv: 2210.15315
handle: 11573/1661239 , 11573/1704623 , 20.500.11850/589368
Cloud computing represents an appealing opportunity for cost-effective deployment of HPC workloads on the best-fitting hardware. However, although cloud and on-premise HPC systems offer similar computational resources, their network architecture and performance may differ significantly. For example, these systems use fundamentally different network transport and routing protocols, which may introduce network noise that can eventually limit the application scaling. This work analyzes network performance, scalability, and cost of running HPC workloads on cloud systems. First, we consider latency, bandwidth, and collective communication patterns in detailed small-scale measurements, and then we simulate network performance at a larger scale. We validate our approach on four popular cloud providers and three on-premise HPC systems, showing that network (and also OS) noise can significantly impact performance and cost both at small and large scale. The full paper of this abstract can be found at https://doi.org/10.1145/3570609.
cloud; HPC; network noise; scalability, Computer Science - Networking and Internet Architecture, Networking and Internet Architecture (cs.NI), Performance (cs.PF), FOS: Computer and information sciences, Computer Science - Performance, Computer Science - Distributed, Parallel, and Cluster Computing, C.4, C.2; C.4, C.2, Distributed, Parallel, and Cluster Computing (cs.DC)
cloud; HPC; network noise; scalability, Computer Science - Networking and Internet Architecture, Networking and Internet Architecture (cs.NI), Performance (cs.PF), FOS: Computer and information sciences, Computer Science - Performance, Computer Science - Distributed, Parallel, and Cluster Computing, C.4, C.2; C.4, C.2, Distributed, Parallel, and Cluster Computing (cs.DC)
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 24 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
| views | 6 | |
| downloads | 26 |

Views provided by UsageCounts
Downloads provided by UsageCounts