
doi: 10.1109/fpl.2011.34
MPI has been used as a parallel programming model for supercomputers and clusters but also in Multiprocessor System-on-Chip. One component of MPI is collective communication and its performance is key for parallel applications to achieve good speedups. Considerable research has been done to optimize such communication by improving the MPI library algorithms. However, these optimizations are focused on the processing nodes (end-points in a network) rather than on the network itself. In this paper, we target a Network-on-Chip (NoC) and modify it to provide hardware support for broadcast and reduce operations for the ArchES-MPI library. This library is a subset implementation of the MPI standard targeting embedded processors and hardware accelerators implemented in FPGAs. The experimental results show that for a system with 24 embedded processors, the broadcast and reduce operations improved up to 11.4-fold and 22-fold, respectively. Higher benefits are expected for larger systems at the expense of a modest increase resource utilization.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 9 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
