
We show how to uniformly distribute data at random (not to be confounded with permutation routing) in two settings that are able to deal with massive data: coarse grained parallelism and external memory. In contrast to previously known work for parallel setups, our method is able to fulfill the three criteria of uniformity, work-optimality and balance among the processors simultaneously. To guarantee the uniformity we investigate the matrix of communication requests between the processors. We show that its distribution is a generalization of the multivariate hypergeometric distribution and we give algorithms to sample it efficiently in the two settings.
uniformly generated communication matrix, Permutations, words, matrices, coarse grained parallelism, F.2.2: Nonnumerical Algorithms and Problems, [INFO.INFO-DS] Computer Science [cs]/Data Structures and Algorithms [cs.DS], Theoretical Computer Science, random shuffling, random permutations, Computational Theory and Mathematics, external memory algorithms, G.2.1.3: Permutations and combinations, Discrete Mathematics and Combinatorics, Analysis of algorithms, Parallel algorithms in computer science
uniformly generated communication matrix, Permutations, words, matrices, coarse grained parallelism, F.2.2: Nonnumerical Algorithms and Problems, [INFO.INFO-DS] Computer Science [cs]/Data Structures and Algorithms [cs.DS], Theoretical Computer Science, random shuffling, random permutations, Computational Theory and Mathematics, external memory algorithms, G.2.1.3: Permutations and combinations, Discrete Mathematics and Combinatorics, Analysis of algorithms, Parallel algorithms in computer science
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 7 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
