
In a prior study (Cioffi & Peroni, 2022), we analysed the available reference extraction tools to understand their performances off-the-shelf – i.e. by using them as they have been configured, without prior training. We evaluated them against a corpus of 56 PDF articles (our gold standard) published in 27 subject areas (Computer Science, Arts and Humanities, Mathematics, etc.). From that analysis, we have identified the two most promising tools for bibliographic reference extraction and parsing, i.e. Anystyle and Grobid, which are CRF based. We have extended such study by training Grobid against an extended gold standard with various training configurations to understand how much the performances improve. As a result, we have also revised the code used for testing and comparing the reference extraction software to make it available also for others to be reused for similar analysis. Othe tests have been performed on OUTCITE and new conversions and evaluations softwares have been created for the purpose. The final aim of this work is to develop a reference extraction service which enables a user to provide a PDF of a scholarly article in input and to have, in return, citation data and bibliographic metadata from all the references that are cited by the given article in a format that enables their ingestion in OpenCitations (Peroni & Shotton, 2020). The demo of the service is available online. This publication is part of my Thesis research for the Digital Humanities and Digital Knowledge Master's Course at University of Bologna Some publications related to the research: The gold standard can be found here, Pagnotta, O. (2024). CEX Project - Dataset and Gold Standard [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10535653. The code can be found here Pagnotta, O. (2024). olgagolgan/CEX-Project: CEX Project Code (software). Zenodo. https://doi.org/10.5281/zenodo.10638757. The output dataset of GROBID, Anystyle and OUTCITE can be found here Pagnotta, O. (2024). CEX Project - Output Dataset (Anystyle, GROBID, OUTCITE) (Version 1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10524898. The training dataset of GROBID can be found here Pagnotta, O. (2024). CEX Project - GROBID annotation aligned Gold Standard (Version 1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10529646. The trained GROBID citation models can be found here Pagnotta, O. (2024). CEX Project - trained GROBID citation models. Zenodo. https://doi.org/10.5281/zenodo.10529709. The final service can be found here Pagnotta, O. and Paolini, L. (2024). opencitations/cec: alpha version (service). Zenodo. https://doi.org/10.5281/zenodo.10635630.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
