Downloads provided by UsageCounts
Minor update from last version: put headers in tab-separated file "distemist_subtrack2_training2_linking.tsv" DisTEMIST corpus DISTEMIST-entities: complete training set (750 clinical cases) DISTEMIST-linking: part 1 of the training set (209 clinical cases) DISTEMIST-linking: part 2 of the training set (375 clinical cases) Introduction The DisTEMIST corpus is a collection of 1000 clinical cases with disease annotations linked with Snomed-CT concepts. All documents are released in the context of the BioASQ DisTEMIST track for CLEF 2022. For more information about the track and its schedule, please visit the website. File structure: The DisTEMIST corpus has been randomly divided into a training set, containing 750 clinical cases, and a test set (584 in the case of subtrack2), consisting of 250 additional cases. Participants must train their systems using the train set and submit predictions for the test set, on which they will be evaluated. The file structure of the corpus is as follows: train_set: text_files: Folder with plain text files of the clinical cases subtrack1_entities: It contains annotations in a tab-separated file (TSV) with the following columns: filename: document name mark: identifier mention id label: mentions type (ENFERMEDAD) off0: starting position of the mention in the document off1: ending position of the mention in the document span: text span subtrack2_linking: It contains annotations in a tab-separated file (TSV) with the following columns: filename: document name mark: identifier mention id label: mentions type (ENFERMEDAD) off0: starting position of the mention in the document off1: ending position of the mention in the document span: text span codes: List of Snomed-CT concept codes linked to the mention. If there is more than one code associated with a mention, they will be concatenated by the symbol “+”. semantic relation: the relationship between the assigned code and the mention. It can be EXACT, when the code corresponds exactly with the mention, or NARROW, when the mention corresponds to a narrower concept than the Snomed-CT code. For instance, the concept “Chorioretinal lacunae” does not exist in Snomed-CT. Then, it is normalized to the Snomed-CT ID 302893000 (“Chorioretinal disorder”). test_set: The 250 clinical cases that will be used to evaluate the systems will be published in accordance with the schedule of the task. Resources Web Evaluation Library DISTEMIST gazetteer DISTEMIST guidelines More resources soon
Funded by the Plan de Impulso de las Tecnologías del Lenguaje (Plan TL).
normalization, NER, gold standard, corpus, Diseases, Entity Linking, NLP, bionlp, Entity Grounding
normalization, NER, gold standard, corpus, Diseases, Entity Linking, NLP, bionlp, Entity Grounding
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 1 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 287 | |
| downloads | 817 |

Views provided by UsageCounts
Downloads provided by UsageCounts