Views provided by UsageCounts
Natural Language Embeddings about Molecular Biology This dataset is concerned with developing a tailored training data set for word and sentence embedding based on biomedical text that has some component associated with molecular work (as opposed to the other range of work indexed in PubMed like non molecular clinical work, studies of human behavior, etc). Raw Data In order to develop natural language embeddings (for words and sentences), we queried PMC and MEDLINE for molecular papers only by using high-level MeSH terms to restrict interest to papers with a molecular focus. We used the following MeSH terms: Cells [A11] Multiprotein Complexes [D05.500] Protein Aggregates [D05.875] Hormones [D06] Enzymes and Coenzymes [D08] Carbohydrates [D08] Lipids [D10] Amino Acids, Peptides and Proteins [D12] Nucleic Acids, Nucleotides and Nucleosides [D13] Biological Factors [D23] Pharmaceutical Preparations [D26] Metabolism [G03] Genetic Phenomena [G06] Queries for these terms use the following string: "cells"[MeSH Terms] OR "Multiprotein Complexes"[mh] OR "Protein Aggregates"[mh] OR "Hormones, Hormone Substitutes, and Hormone Antagonists"[mh] OR "Enzymes and Coenzymes"[mh] OR "Carbohydrates"[mh] OR "Lipids"[mh] OR "Amino Acids, Peptides, and Proteins"[mh] OR "Nucleic Acids, Nucleotides, and Nucleosides"[mh] OR "Biological Factors"[mh] OR "Pharmaceutical Preparations"[mh] OR "Metabolism"[mh] OR "Cell Physiological Phenomena"[mh] OR "Genetic Phenomena"[mh] PubMed returns 11,447,521 abstracts. PMC returns, 1,720,266 documents, 509,722 of these are open access. We downloaded, parsed and concatenated 403,825 PMC open access documents into a single file `molecular_oa_pmc.tsv`. This is a 33GB TSV file with the following columns: File:Paragraph - a unique identifier for each paragraph SentenceId - the local number of the sentence in the document Sentence Text - tokenized text of the sentence (based on ClearTk's TokenAnnotator.java) Codes - exLink for the presence of a citation, inLink for the presence of link to a Figure Figures - Figure codes Headings - High level section of the paper Offset_Begin - offset of the start of the sentence within the paper Offset_End - offset of the start of the sentence within the paper We repeated the same process for PubMed abstracts to generate a 3.6G file (`molecular_oa_medline.tsv`) with three columns: Pubmed ID A Boolean value indicating whether the article is a review Text We concatenated the text columns of these two files into a single 30GB file (`molecular_oa.txt`) where each line is a single sentence and the text is fully tokenized. These three files are archived in `molecular_oa_raw_text.tar.gz`. Fasttext Embedding We trained a fasttext model on the raw training data (https://fasttext.cc/) using the standard `skipgram` parameter. A gzipped copy of the word embeddings is included in `fasttext.model.vec.gz`
This work was funded under DARPA Big Mechanism program under ARO contract W911NF-14-1-0436 and NIH grant 1 R01 LM012592 ('EVIDENCE EXTRACTION SYSTEMS FOR THE MOLECULAR INTERACTION LITERATURE')
Molecular Biology, NL Embedding
Molecular Biology, NL Embedding
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 8 |

Views provided by UsageCounts