Downloads provided by UsageCounts
Open repositories are open to EVERYONE. Unfortunately, that includes machines, or robots / “bots”, roaming around open repositories for various reasons, or even by chance, thus creating non-human-generated statistics for the repositories they visit. This issue is especially significant to open repositories, as subscription-based databases usually place more obstacles in the path of “bots”, including authentication and authorization. While acceptable usage of open repositories encompasses any human interaction meant to satisfy intellectual curiosity or engage professionally with the repository’s content, machinegenerated engagement is usually in bulk and does not serve research goals that the repository is designed to assist with. Machine-generated usage statistics could skew the accuracy of impact and engagement metrics for individual publications or entire repositories, thus impacting funding, perception, content acquisition policy and more. Therefore, an accurate measuring of engagement, usage, and their various metrics for open repositories requires effective methods of weeding out non-human-generated statistics. We have come up with a workflow to minimize machine-generated statistics as much as we can, identifying seven “layers” of “bot” activity, ranked by ease of detection and codified by the colors of the rainbow (ROYGBIV). The rainbow-theme color coding strives to divide machine generated activity in a repository by the ease of detection by information specialists administrating an open repository, suggesting methods and tools for the detection and handling of each category, when applicable. Red: DDoS attacks Orange: Crawler HTTP agents Yellow: Known blacklisted IP addresses Green: Newly found blacklisted IP addresses Blue: Non-blacklisted IP addresses behaving suspiciously Indigo: Bot disguised as VPN Violet: Bot client arrays Open repository statistics can virtually never be 100% “clean” of machine-generated impact. However, adhering to the workflow and using freely available tools can significantly reduce “noise” in the measurement of repository metrics, helping publishers and organizations make better informed decisions and providing a realistic usage picture. In this poster session we will discuss the various types of machine-generated statistics indicators and how to best detect them. The poster will also include future trends in bot activity and detection.
doi:10.5281/ZENODO.6367305
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 2 | |
| downloads | 8 |

Views provided by UsageCounts
Downloads provided by UsageCounts