research data . Dataset . Under curation

bioRxiv 10k

Daniel Ecer;
Open Access English
  • Publisher: Zenodo
This dataset is a CC-BY 4.0 subset of what bioRxiv kindly made available: It is randomized and split into train (6,000), validation (2,000) and test (2,000) subsets - 10,000 PDF / XML pairs in total. The zip files further contain file lists of smaller subsets that used the subject area to potentially create a balanced subset. The zip is similar in structure to the "PMC sample 1943" dataset that was created as part of: (a working link is available from: Therefore it is well suited for evaluation of PDF to XML conversio...
Persistent Identifiers
Download from
Any information missing or wrong?Report an Issue