
This paper proposes ACC-AMG, an optimized AMG framework that further improves the efficiency of sparse general matrix-matrix multiplication (SpGEMM) on modern GPUs. Considering that SpGEMM is a major bottleneck in the AMG setup phase, we propose a performance-model-based method that takes the multi-dimensional features of two matrix blocks as input and maps each block-level computation to the most suitable computing unit, i.e., Tensor Cores (TCs) or CUDA Cores (CCs). To better exploit this adaptive execution strategy, ACC-AMG further incorporates targeted register-pressure reduction for CC execution, a block-wise load-balancing scheme for TC execution, and mixed-precision support. To better utilize algorithm components in existing libraries, ACC-AMG is integrated into the hypre framework.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
