Tokenization for Molecular Foundation Models

descriptionPublicationkeyboard_double_arrow_right Article , Other literature type , Preprint 23 Jan 2026Embargo end date: 01 Jan 2024 English Publisher:American Chemical Society (ACS)Journal:Journal of Chemical Information and Modeling, volume 66, pages 1,384-1,393 (issn: 1549-9596, eissn: 1549-960X,

Copyright policy )

Authors: Wadell, Alexius; Bhutani, Anoushka; Viswanathan, Venkatasubramanian;

doi: 10.1021/acs.jcim.5c01856 , 10.5281/zenodo.13761263 , 10.5281/zenodo.13761262 , 10.48550/arxiv.2409.15370 , 10.5281/zenodo.17463868 , 10.5281/zenodo.17835329

pmid: 41575906

arXiv: 2409.15370

Tokenization for Molecular Foundation Models

- Summary
- Subjects
- Metrics

Abstract

Data Drop and Source Code for Tokenization for Molecular Foundation Models File Contents smirk-0.1.1.tar.gz Source code for the Smirk Tokenizer TokenizerStats.tar.gz Source code for paper plots, tokenizer analysis, substring ambiguities, and fine-tuned models. ngram_tokenizer_stats.tar.xz N-gram models and tabulated Cross-Entropy and Information Loss for all evaluated tokenizers serialized using JLD2. The source code for working with and loading these files is in `TokenizerStats.tar.gz`. Decompresses to ~32.4 GiB. Summary statistics, fixed-effects models, and coverage statistics are additionally provided. tmQM.tar.qz Source code and generated property prediction dataset constructed from the tmQM dataset. The dataset is provided in the Apache Arrow format and provides molecules transcoded into OpenSMILES from the source XYZ files. The dataset is readily readable using HuggingFace's load_dataset function. safetensors-models.tar.xz Safetensor checkpoints for all pre-trained and fine-tuned models (270) trained as part of this work. Instructions for loading these models are provided in TokenizerStats.tar.gz. Decompresses to ~31.7 GiB. pubchem_ambiguous_substrings.csv.xz Collection of molecules with ambiguous substrings (i.e., Sc, Cn, Sn, etc.) retrieved from PubChem. Generation details are provided in the paper's supporting information. The filtering script (check_ambi_smiles.py) is included within TokenizerStats.tar.gz

Related Organizations

University of Michigan–Flint
United States
University of Michigan
United States
Carnegie Mellon University
United States
University of Michigan–Ann Arbor
United States

Keywords

Machine Learning, Biomolecules, Chemical Physics (physics.chem-ph), FOS: Computer and information sciences, Chemical Physics, Artificial Intelligence (cs.AI), Artificial Intelligence, FOS: Biological sciences, FOS: Physical sciences, Biomolecules (q-bio.BM), Machine Learning (cs.LG)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	1
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

1

Top 10%

Average

Green