Downloads provided by UsageCounts
Here are the datasets used for GROBID end-to-end benchmarking covering: - metadata extraction, - bibliographical reference extraction, parsing and citation context identification, and - full text body structuring. The following collections are included: - a PubMedCentral gold-standard dataset called PMC_sample_1943, compiled by Alexandru Constantin. The dataset, around 1.5GB in size, contains 1943 articles from 1943 different journals corresponding to the publications from a 2011 snapshot. For each article, we have a PDF file and a NLM XML file. - a bioRxiv dataset called biorxiv-10k-test-2000 of 2000 preprint articles originally compiled with care and published by Daniel Ecer, available on Zenodo. The dataset contains for each article a PDF file and the corresponding reference NLM file (manually created by bioRxiv). The NLM files have been further systematically reviewed and annotated with additional markup corresponding to data and code availability statements and funding statements by the Grobid team. Around 5.4G in size. - a set of 1000 PLOS articles, called PLOS_1000, randomly selected from the full PLOS Open Access collection. Again, for each article, the published PDF is available with the corresponding publisher JATS XML file, around 1.3GB total size. - a set of 984 articles from eLife, called eLife_984, randomly selected from their open collection available on GitHub. Every articles come with the published PDF, the publisher JATS XML file and the eLife public HTML file (as bonus, not used), all in their latest version, around 4.5G total. For each of these datasets, the directory structure is the same and documented here. Further information on Grobid benchmarking and how to run it: https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/. Latest benchmarking scores are also available in the Grobid documentation: https://grobid.readthedocs.io/en/latest/Benchmarking/ These resources are originally published under CC-BY license. Our additional annotations are similarly under CC-BY. We thank NIH, bioRxiv, PLOS and eLife for making these resources Open Access and reusable.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 117 | |
| downloads | 50 |

Views provided by UsageCounts
Downloads provided by UsageCounts