
arXiv: 2008.03073
AbstractThe power law is useful in describing count phenomena such as network degrees and word frequencies. With a single parameter, it captures the main feature that the frequencies are linear on the log‐log scale. Nevertheless, there have been criticisms of the power law, for example, that a threshold needs to be preselected without its uncertainty quantified, that the power law is simply inadequate, and that subsequent hypothesis tests are required to determine whether the data could have come from the power law. We propose a modeling framework that combines two different generalizations of the power law, namely the generalized Pareto distribution and the Zipf‐polylog distribution, to resolve these issues. The proposed mixture distributions are shown to fit the data well and quantify the threshold uncertainty in a natural way. A model selection step embedded in the Bayesian inference algorithm further answers the question whether the power law is adequate.
polyalgorithm, Social and Information Networks (cs.SI), FOS: Computer and information sciences, Parametric inference, 330, Computer Science - Social and Information Networks, Nonparametric inference, Applications of statistics, degree distribution, Statistics - Applications, Methodology (stat.ME), Markov chain Monte Carlo, generalized Pareto, Applications (stat.AP), Bayesian model selection, threshold uncertainty, Statistics - Methodology
polyalgorithm, Social and Information Networks (cs.SI), FOS: Computer and information sciences, Parametric inference, 330, Computer Science - Social and Information Networks, Nonparametric inference, Applications of statistics, degree distribution, Statistics - Applications, Methodology (stat.ME), Markov chain Monte Carlo, generalized Pareto, Applications (stat.AP), Bayesian model selection, threshold uncertainty, Statistics - Methodology
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 5 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
