
Data Drop and Source Code for Tokenization for Molecular Foundation Models File Contents smirk-0.1.1.tar.gz Source code for the Smirk Tokenizer TokenizerStats.tar.gz Source code for paper plots, tokenizer analysis, substring ambiguities, and fine-tuned models. ngram_tokenizer_stats.tar.xz N-gram models and tabulated Cross-Entropy and Information Loss for all evaluated tokenizers serialized using JLD2. The source code for working with and loading these files is in `TokenizerStats.tar.gz`. Decompresses to ~32.4 GiB. Summary statistics, fixed-effects models, and coverage statistics are additionally provided. tmQM.tar.qz Source code and generated property prediction dataset constructed from the tmQM dataset. The dataset is provided in the Apache Arrow format and provides molecules transcoded into OpenSMILES from the source XYZ files. The dataset is readily readable using HuggingFace's load_dataset function. safetensors-models.tar.xz Safetensor checkpoints for all pre-trained and fine-tuned models (270) trained as part of this work. Instructions for loading these models are provided in TokenizerStats.tar.gz. Decompresses to ~31.7 GiB. pubchem_ambiguous_substrings.csv.xz Collection of molecules with ambiguous substrings (i.e., Sc, Cn, Sn, etc.) retrieved from PubChem. Generation details are provided in the paper's supporting information. The filtering script (check_ambi_smiles.py) is included within TokenizerStats.tar.gz
Machine Learning, Biomolecules, Chemical Physics (physics.chem-ph), FOS: Computer and information sciences, Chemical Physics, Artificial Intelligence (cs.AI), Artificial Intelligence, FOS: Biological sciences, FOS: Physical sciences, Biomolecules (q-bio.BM), Machine Learning (cs.LG)
Machine Learning, Biomolecules, Chemical Physics (physics.chem-ph), FOS: Computer and information sciences, Chemical Physics, Artificial Intelligence (cs.AI), Artificial Intelligence, FOS: Biological sciences, FOS: Physical sciences, Biomolecules (q-bio.BM), Machine Learning (cs.LG)
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 1 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
