
doi: 10.1002/cplx.20022
sug-gests the presence of linguistic fea-tures in eukaryotic genomes, and thatthis finding “may lead to importantclues about the evolution of lan-guages, DNA, and information pro-cessing.” We believe that these conclu-sions are too far-fetched, and that thepaper attaches Zipf’s law too closely tolanguage alone, as if it is a feature ex-clusive to human languages. In addi-tion, we would like to mention that (i)the power-law decay of the number ofoccurrences of protein domains withtheir rank has been found for dozensof different organisms in Refs. 2 and 3and that (ii) the ubiquitous presenceof those power laws could be ex-plained by simple models of sequenceevolution.Admittedly, Zipf’s law was first dis-covered in English, German, and otherhuman languages, but George Kings-ley Zipf himself explored the similarplot in city population, settlementsize, and human migration [4]. Today,Zipf’s law is known to appear in manycontexts unrelated to languages: indi-vidual incomes, company sizes, inter-net traffic, web page popularities,number of citations, microarray andgene expression data [5–7], etc. (For anattempt to summarize various obser-vations of Zipf’s law, see e.g., [8]). Evenrandomly generated texts, treatingcharacter strings between two blankspaces as “words,” exhibit a rank-fre-quency statistics consistent with Zipf’slaw [9].Can we say that these examples ofZipf’s law exhibit “linguistic features?”Clearly not. The size of an earthquakehas nothing to do with linguistics, andthe text typed by a monkey is not thework of Shakespeare. Yet, both textsand the size distribution of earthquakes are consistent with Zipf’s law.Using the language metaphor has along history in molecular biology, andone of its earliest fruits may be thediscovery of the genetic code [10].However, as powerful as it might be asa metaphor, Zipf’s law observed ingenomic data is more likely caused bysimple processes of sequence evolu-tion mentioned in [2, 3] than by forcesthat attempt to endow biomoleculeswith linguistic features.Both Refs. 2 and 3 show that ran-domly occurring mutations and dupli-cations are sufficient to generate apower-law distribution of gene familysizes. This distribution is not identi-cal—but related—to the rank-fre-quency plot. Specifically, there is aone-to-one relationship between thesize distribution and the rank-fre-quency plot (see, e.g., [8, 11–14]). Al-fred Lotka uses the size distribution inhis study of scientific productivity [15],and George Miller emphasized thatthe size distribution is also part of“Zipf’s curve” [16].The one-to-one relationship be-tween Zipf’s law and Lotka’s law is easyto understand. According to Zipf’s law,the frequency
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
