
doi: 10.3233/sw-200374
While workflow systems have improved the repeatability of scientific experiments, the value of the processed (intermediate) data have been overlooked so far. In this paper, we argue that the intermediate data products of workflow executions should be seen as first-class objects that need to be curated and published. Not only will this be exploited to save time and resources needed when re-executing workflows, but more importantly, it will improve the reuse of data products by the same or peer scientists in the context of new hypotheses and experiments. To assist curator in annotating (intermediate) workflow data, we exploit in this work multiple sources of information, namely: (i) the provenance information captured by the workflow system, and (ii) domain annotations that are provided by tools registries, such as Bio.Tools. Furthermore, we show, on a concrete bioinformatics scenario, how summarising techniques can be used to reduce the machine-generated provenance information of such data products into concise human- and machine-readable annotations.
020, provenance, scientific workflows, bioinformatics, [INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI], 004, Linked Data, [INFO.INFO-BI]Computer Science [cs]/Bioinformatics [q-bio.QM], 005.7, Organisation des données, data summaries, FAIR
020, provenance, scientific workflows, bioinformatics, [INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI], 004, Linked Data, [INFO.INFO-BI]Computer Science [cs]/Bioinformatics [q-bio.QM], 005.7, Organisation des données, data summaries, FAIR
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 4 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
