Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Conference object . 2016
License: CC BY
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Other literature type . 2016
License: CC BY
Data sources: ZENODO
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Conference object . 2016
License: CC BY
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Other literature type . 2016
License: CC BY
Data sources: ZENODO
versions View all 2 versions
addClaim

A Machine Learning Approach For Data Source And Type Identification To Support Metadata Discovery

Authors: Wen, Jingran; Gouripeddi, Ramkiran; Facelli, Julio;

A Machine Learning Approach For Data Source And Type Identification To Support Metadata Discovery

Abstract

Current approaches to metadata discovery are dependent on manual curations which are time consuming processes and it is critical to develop automatic and/or semiautomatic metadata discovery methods to realize the full potential of Big Data technologies in biomedicine, enhance research reproducibility and increase efficiency in translational biomedical sciences. We are developing a two-step metadata discovery workflow: (1) Identification of data source and type using their intrinsic document structure, and (2) Discovery of detailed metadata within the data by associating specific metadata discovery tools based on the data’s source and type ascertained from (1). Here we discuss our initial results for (1) using machine learning. In this work we included biomedical data from various sources and in different file formats: 1) Human genetic variants - ClinVar (XML, tab-delimited), 2) Protein structure (PDB, mmCIF, XML) 3) Biomedical literature, and General English corpus. We tokenized the data files using the Natural Language Toolkit and considered various document structural features: normalized count of numerical tokens, negative numerical tokens, number of words, number of capitalized words, number of words with all upper letters, and median length of tokens. We developed a decision tree model with these structural features to classify these data and types, and evaluated the performance of the model using 10-fold cross-validation and a test set for the following metrics: precision, recall, and F1 score using scikit-learn. Our model was able to distinguish protein structure, genetic variant, scientific paper and general English files with an average F1 score of 0.997, 0.997, 0.886 and 0.919 when evaluated using cross-validation, and 1, 0.999, 0.980, and 0.935 when using independent test sets.. Our approach shows it is possible to automatically identify data sources and types using only document structural features and therefore reasonable to programmatically associate metadata extraction tools specific for each data source and type as next steps.

Related Organizations
Keywords

BD2K_AHM, metadata discovery, machine learning, data source classification

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
    OpenAIRE UsageCounts
    Usage byUsageCounts
    visibility views 10
    download downloads 1
  • 10
    views
    1
    downloads
    Powered byOpenAIRE UsageCounts
Powered by OpenAIRE graph
Found an issue? Give us feedback
visibility
download
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
views
OpenAIRE UsageCountsViews provided by UsageCounts
downloads
OpenAIRE UsageCountsDownloads provided by UsageCounts
0
Average
Average
Average
10
1
Green