Identity resolution of software metadata using Large Language Models

Name: Identity resolution of software metadata using Large Language Models
Keywords: Software Engineering (cs.SE), FOS: Computer and information sciences, Digital Libraries, Software Engineering, Digital Libraries (cs.DL), Computation and Language, Computation and Language (cs.CL)

Martin del Pico, Eva; Capella-Gutierrez, Salvador; Gelpí, Josep Lluís

Found an issue? Give us feedback

arXiv.org e-Print Ar...arrow_drop_down

arXiv.org e-Print Archive

Preprint . 2025

Data sources: arXiv.org e-Print Archive

ZENODO

Preprint . 2025

License: CC BY

Data sources: Datacite

ZENODO

Preprint . 2025

License: CC BY

Data sources: Datacite

https://dx.doi.org/10.48550/ar...

Article . 2025

License: CC BY NC ND

Data sources: Datacite

DBLP

Article

Data sources: DBLP

Identity resolution of software metadata using Large Language Models

descriptionPublicationkeyboard_double_arrow_right Preprint , Article 01 Jan 2025Embargo end date: 01 Jan 2025Publisher:ZenodoJournal:CoRR, volume abs/2505.23500

Authors: Martin del Pico, Eva; Capella-Gutierrez, Salvador; Gelpí, Josep Lluís;

doi: 10.5281/zenodo.15546632 , 10.5281/zenodo.15546631 , 10.48550/arxiv.2505.23500

arXiv: 2505.23500

Identity resolution of software metadata using Large Language Models

- Summary
- Subjects
- Metrics

Abstract

Software is an essential component of research. However, little attention has been paid to it compared with that paid to research data. Recently, there has been an increase in efforts to acknowledge and highlight the importance of software in research activities. Structured metadata from platforms like bio.tools, Bioconductor, and Galaxy ToolShed offers valuable insights into research software in the Life Sciences. Although originally intended to support discovery and integration, this metadata can be repurposed for large-scale analysis of software practices. However, its quality and completeness vary across platforms, reflecting diverse documentation practices. To gain a comprehensive view of software development and sustainability, consolidating this metadata is necessary—but requires robust mechanisms to address its heterogeneity and scale. This article presents an evaluation of instruction-tuned large language models for the task of software metadata identity resolution—a critical step in assembling a cohesive collection of research software. Such a collection is the reference component for the Software Observatory at OpenEBench, a platform that aggregates metadata to monitor the "FAIRness" of research software in the Life Sciences. We benchmarked multiple models against a human-annotated gold standard, examined their behavior on ambiguous cases, and introduced an agreement-based proxy for high-confidence automated decisions. The proxy achieved high precision and statistical robustness, while also highlighting the limitations of current models and the broader challenges of automating semantic judgment in FAIR-aligned software metadata across registries and repositories.

Related Organizations

University of Barcelona
Spain
Barcelona Supercomputing Center
Spain

Keywords

Software Engineering (cs.SE), FOS: Computer and information sciences, Digital Libraries, Software Engineering, Digital Libraries (cs.DL), Computation and Language, Computation and Language (cs.CL)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

0

Average

Green