"Why do you cite?" An investigation on Citation Intents and Decision-Making Classification Processes

descriptionPublicationkeyboard_double_arrow_right Other literature type , Article 16 Jun 2024 English Publisher:ZenodoJournal:CoRR, volume abs/2407.13329

Authors: Paolini, Lorenzo; Vahdati, Sahar; Di Iorio, Angelo; Wardenga, Robert; Heibi, Ivan; Peroni, Silvio;

doi: 10.5281/zenodo.11841799

"Why do you cite?" An investigation on Citation Intents and Decision-Making Classification Processes

- Summary
- Metrics

Abstract

Abstract Identifying the reason for which an author cites another work is essential to understand the nature of scientific contributions and to assess their impact. Citations are one of the pillars of scholarly communication and most metrics employed to analyze these conceptual links are based on quantitative observations. Behind the act of referencing another scholarly work there is a whole world of meanings and needs that needs to be proficiently and effectively revealed. This study emphasizes the importance of trustfully classifying citation intents to provide more comprehensive and insightful analyses in research assessment. We address this task by presenting a study utilizing advanced Ensemble Strategies for Citation Intent Classification (CIC) incorporating Language Models (LMs) and employing Explainable AI (XAI) techniques to enhance the interpretability and trustworthiness of models’ predictions. Our approach involves two ensemble classifiers that utilize fine-tuned SciBERT and XLNet models as baselines. We further demonstrate the critical role of section titles as a feature in improving models’ performances. The study also introduces a web application developed with Flask and currently available at http://137.204.64.4:81/cic/classifier, aimed at classifying citation intents. One of our models sets as a new state-of-the-art (SOTA) with an 89.46% Macro-F1 score on the SciCite bench- mark. The integration of SHAP and LIME for explainability provides insights into the decision-making processes, highlighting the contributions of individual words for the binary predictions produced by the base models of the ensemble, and highlighting the role that base models have in the final classification of a sentence, performed by a meta-classifier head. The findings suggest that the inclusion of section titles significantly enhances classification performances in the CIC task. Our contributions provide useful insights for developing more robust datasets and methodologies, thus fostering a deeper understanding of scholarly communication.

Related Organizations

TU Dresden
Germany
Alma Mater Studiorum University of Bologna
Italy
OpenCitations
Italy
Institute of Computer Vision and Applied Computer Sciences
Germany

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average