
usiGrabber is a scalable framework for assembling large and diverse mass-spectrometry datasets ready to be used for machine learning use cases As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra and used it to retrain a binary phosphorylation classifier. This dataset and the corresponding model weights are available in this record. The publication also includes the complete database, which contains spectrum information and metadata for over 800 million spectra present in the PRIDE database. Because of its size, it had to be split into multiple uploads. In order to reconstruct the entire database, you must download all related records. Once you have downloaded all records, extract the archives and refer to usiGrabber - db_export for instructions for reassembly. Related records: peptide_spectrum_matches table: https://zenodo.org/records/18890370 psm_peptide_evidence table: https://zenodo.org/records/18864164 Other, smaller tables: https://zenodo.org/records/18873214
Proteomics, Machine Learning, Deep Learning, Mass Spectrometry, Data Curation
Proteomics, Machine Learning, Deep Learning, Mass Spectrometry, Data Curation
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
