
This data set pertains to the largest works in the OpenITI corpus at or prior to 1000 AH and is based on the 2023.1.8 release of the corpus and the corresponding text reuse data between the books in the corpus, which is generated by running using passim on the corpus. We wanted to understand the extent to which a small number of persons produced a substantial percentage of the OpenITI corpus, on a word-count basis. We call the authors with work(s) over a million words the ‘millionaires’. The data will be analysed in forthcoming publications by the KITAB project team, including a monograph by Sarah Bowen Savant under contract with Edinburgh University Press. KITAB is funded by the European Research Council under the European Union’s Horizon 2020 research and innovation programme, awarded to the KITAB project (Grant Agreement No. 772989, PI Sarah Bowen Savant), hosted at Aga Khan University, London. In addition, it has received funding from the Qatar National Library to aid in the adaptation of the passim algorithm for Arabic. KITAB’s text reuse data is published on Zenodo and each version is the output of a separate run. The version number of each release corresponds to the corpus releases.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
